ScholarGate
Asystent

Porównaj metody

Przeglądaj wybrane metody obok siebie; wiersze, które się różnią, są wyróżnione.

Wielojęzyczny Transformer Wizyjny×Multimodal Vision Transformer×
DziedzinaUczenie głębokieUczenie głębokie
RodzinaMachine learningMachine learning
Rok powstania2021–20232021
TwórcaDosovitskiy et al. (ViT base); multilingual extension by multiple groups (2021–2023)Dosovitskiy et al. (ViT); Radford et al. (CLIP multimodal ViT)
TypTransformer-based vision model with multilingual capabilitiesMultimodal transformer model
Źródło pierwotneDosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations (ICLR 2021). link ↗Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR). link ↗
Inne nazwyMultilingual ViT, Cross-lingual Vision Transformer, Multilingual Visual Transformer, ML-ViTMultimodal ViT, vision-language transformer, cross-modal vision transformer, multi-modal ViT
Pokrewne45
PodsumowanieMultilingual Vision Transformer (Multilingual ViT) extends the Vision Transformer architecture to operate across multiple languages, enabling image understanding and image-text reasoning in multilingual or cross-lingual settings. It combines patch-based image encoding with multilingual text representations, allowing a single model to serve diverse linguistic communities for tasks such as image captioning, visual question answering, and cross-lingual image retrieval.Multimodal Vision Transformer (Multimodal ViT) extends the Vision Transformer architecture to jointly process and align representations from multiple modalities — typically images and text — using self-attention and cross-attention mechanisms. By learning shared or aligned embedding spaces across modalities, it enables tasks such as visual question answering, image-text retrieval, visual grounding, and image captioning.
ScholarGateZbiór danych
  1. v1
  2. 2 Źródła
  3. PUBLISHED
  1. v1
  2. 2 Źródła
  3. PUBLISHED

Przejdź do wyszukiwania Pobierz slajdy

ScholarGatePorównaj metody: Multilingual vision transformer · Multimodal Vision Transformer. Pobrano 2026-06-18 z https://scholargate.app/pl/compare