ScholarGate
Asistent

Porovnat metody

Prohlédněte si vybrané metody vedle sebe; řádky, které se liší, jsou zvýrazněny.

Vícejazyčný Vision Transformer×Multimodální Vision Transformer×
OborHluboké učeníHluboké učení
RodinaMachine learningMachine learning
Rok vzniku2021–20232021
TvůrceDosovitskiy et al. (ViT base); multilingual extension by multiple groups (2021–2023)Dosovitskiy et al. (ViT); Radford et al. (CLIP multimodal ViT)
TypTransformer-based vision model with multilingual capabilitiesMultimodal transformer model
Původní zdrojDosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. International Conference on Learning Representations (ICLR 2021). link ↗Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations (ICLR). link ↗
Další názvyMultilingual ViT, Cross-lingual Vision Transformer, Multilingual Visual Transformer, ML-ViTMultimodal ViT, vision-language transformer, cross-modal vision transformer, multi-modal ViT
Příbuzné45
ShrnutíMultilingual Vision Transformer (Multilingual ViT) extends the Vision Transformer architecture to operate across multiple languages, enabling image understanding and image-text reasoning in multilingual or cross-lingual settings. It combines patch-based image encoding with multilingual text representations, allowing a single model to serve diverse linguistic communities for tasks such as image captioning, visual question answering, and cross-lingual image retrieval.Multimodal Vision Transformer (Multimodal ViT) extends the Vision Transformer architecture to jointly process and align representations from multiple modalities — typically images and text — using self-attention and cross-attention mechanisms. By learning shared or aligned embedding spaces across modalities, it enables tasks such as visual question answering, image-text retrieval, visual grounding, and image captioning.
ScholarGateDatová sada
  1. v1
  2. 2 Zdroje
  3. PUBLISHED
  1. v1
  2. 2 Zdroje
  3. PUBLISHED

Přejít na hledání Stáhnout prezentaci

ScholarGatePorovnat metody: Multilingual vision transformer · Multimodal Vision Transformer. Získáno 2026-06-18 z https://scholargate.app/cs/compare