ScholarGate
Asistent

Porovnať metódy

Prezrite si vybrané metódy vedľa seba; riadky, ktoré sa líšia, sú zvýraznené.

Multimodálny Doc2Vec×Multimodálny Transformer×
OdborHlboké učenieHlboké učenie
RodinaMachine learningMachine learning
Rok vzniku2014–20172019–2021
TvorcaLe, Q. V. & Mikolov, T. (Doc2Vec core); multimodal extensions by various authors post-2014Lu et al. (ViLBERT); Radford et al. (CLIP)
TypMultimodal document embeddingCross-modal attention-based deep learning model
Pôvodný zdrojLe, Q. V., & Mikolov, T. (2014). Distributed Representations of Sentences and Documents. Proceedings of the 31st International Conference on Machine Learning (ICML), PMLR 32(2), 1188–1196. link ↗Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Advances in Neural Information Processing Systems (NeurIPS), 32. link ↗
Ďalšie názvyMultimodal Paragraph Vector, Cross-modal Doc2Vec, Multi-source PV-DM, Multimodal Document Embeddingmultimodal attention model, cross-modal transformer, vision-language transformer, multi-modal fusion transformer
Príbuzné65
ZhrnutieMultimodal Doc2Vec extends the Doc2Vec paragraph-vector framework to incorporate information from more than one modality — typically text alongside images, audio, or structured metadata — producing a shared document-level embedding that captures semantics from multiple sources simultaneously. It is used for cross-modal retrieval, multi-source classification, and document representation where text alone is insufficient.A Multimodal Transformer extends the standard Transformer architecture to process and jointly reason over two or more input modalities — most commonly text and images, but also audio, video, or structured data. Cross-modal attention layers allow information from one modality to inform representations in another, enabling tasks such as visual question answering, image captioning, and multimodal sentiment analysis.
ScholarGateDátová sada
  1. v1
  2. 2 Zdroje
  3. PUBLISHED
  1. v1
  2. 2 Zdroje
  3. PUBLISHED

Prejsť na hľadanie Stiahnuť snímky

ScholarGatePorovnať metódy: Multimodal Doc2Vec · Multimodal Transformer. Získané 2026-06-17 z https://scholargate.app/sk/compare