ScholarGate
Asisten

Bandingkan metode

Tinjau metode pilihan Anda berdampingan; baris yang berbeda akan disorot.

Multimodal Doc2Vec×Transformer Multimodal×
BidangPembelajaran MendalamPembelajaran Mendalam
KeluargaMachine learningMachine learning
Tahun asal2014–20172019–2021
PencetusLe, Q. V. & Mikolov, T. (Doc2Vec core); multimodal extensions by various authors post-2014Lu et al. (ViLBERT); Radford et al. (CLIP)
TipeMultimodal document embeddingCross-modal attention-based deep learning model
Sumber perintisLe, Q. V., & Mikolov, T. (2014). Distributed Representations of Sentences and Documents. Proceedings of the 31st International Conference on Machine Learning (ICML), PMLR 32(2), 1188–1196. link ↗Lu, J., Batra, D., Parikh, D., & Lee, S. (2019). ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Advances in Neural Information Processing Systems (NeurIPS), 32. link ↗
AliasMultimodal Paragraph Vector, Cross-modal Doc2Vec, Multi-source PV-DM, Multimodal Document Embeddingmultimodal attention model, cross-modal transformer, vision-language transformer, multi-modal fusion transformer
Terkait65
RingkasanMultimodal Doc2Vec extends the Doc2Vec paragraph-vector framework to incorporate information from more than one modality — typically text alongside images, audio, or structured metadata — producing a shared document-level embedding that captures semantics from multiple sources simultaneously. It is used for cross-modal retrieval, multi-source classification, and document representation where text alone is insufficient.A Multimodal Transformer extends the standard Transformer architecture to process and jointly reason over two or more input modalities — most commonly text and images, but also audio, video, or structured data. Cross-modal attention layers allow information from one modality to inform representations in another, enabling tasks such as visual question answering, image captioning, and multimodal sentiment analysis.
ScholarGateSet data
  1. v1
  2. 2 Sumber
  3. PUBLISHED
  1. v1
  2. 2 Sumber
  3. PUBLISHED

Ke halaman pencarian Unduh salindia

ScholarGateBandingkan metode: Multimodal Doc2Vec · Multimodal Transformer. Diakses 2026-06-17 dari https://scholargate.app/id/compare