Sammenlign metoder
Gjennomgå de valgte metodene side om side; rader som avviker, er uthevet.
| Fler-språklig emnemodellering× | Multilingual Transformer× | |
|---|---|---|
| Fagfelt | Dyp læring | Dyp læring |
| Familie | Machine learning | Machine learning |
| Opprinnelsesår≠ | 2009 | 2019–2020 |
| Opphavsperson≠ | Mimno, D., Wallach, H. M., et al. | Devlin et al. (mBERT); Conneau et al. (XLM-R) |
| Type≠ | Probabilistic topic model (multilingual extension) | Pre-trained cross-lingual language model |
| Opprinnelig kilde≠ | Mimno, D., Wallach, H. M., Naradowsky, J., Smith, D. A., & McCallum, A. (2009). Polylingual topic models. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 880–889. ACL. link ↗ | Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of NAACL-HLT 2019, pp. 4171–4186. Association for Computational Linguistics. DOI ↗ |
| Alias | cross-lingual topic model, polylingual LDA, multilingual LDA, MLTM | multilingual LM, cross-lingual transformer, mBERT-style model, multilingual pre-trained model |
| Relaterte≠ | 5 | 4 |
| Sammendrag≠ | Multilingual topic modeling extends probabilistic topic models such as LDA to corpora spanning two or more languages, inferring shared latent topics across language boundaries. By tying topic distributions across languages, it enables cross-lingual document analysis, comparable topic discovery, and information retrieval without requiring full parallel corpora. | A multilingual transformer is a pre-trained language model built on the transformer architecture and trained jointly on text from dozens to over one hundred languages. Models such as mBERT and XLM-RoBERTa learn shared cross-lingual representations, enabling zero-shot or few-shot transfer: a model fine-tuned on English data can often be applied directly to French, German, Arabic, or Chinese without language-specific labels. |
| ScholarGateDatasett ↗ |
|
|