So sánh phương pháp
Xem các phương pháp đã chọn cạnh nhau; những hàng khác biệt được làm nổi bật.
| Mô hình hóa chủ đề đa ngôn ngữ× | Nhúng câu đa ngôn ngữ× | |
|---|---|---|
| Lĩnh vực | Học sâu | Học sâu |
| Họ | Machine learning | Machine learning |
| Năm ra đời≠ | 2009 | 2019–2022 |
| Người khởi xướng≠ | Mimno, D., Wallach, H. M., et al. | Reimers, N. & Gurevych, I.; Feng, F. et al. (Google) |
| Loại≠ | Probabilistic topic model (multilingual extension) | Cross-lingual representation learning |
| Công trình gốc≠ | Mimno, D., Wallach, H. M., Naradowsky, J., Smith, D. A., & McCallum, A. (2009). Polylingual topic models. In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 880–889. ACL. link ↗ | Reimers, N. & Gurevych, I. (2020). Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation. Proceedings of EMNLP 2020, 4512–4525. link ↗ |
| Tên gọi khác | cross-lingual topic model, polylingual LDA, multilingual LDA, MLTM | multilingual sentence representations, cross-lingual sentence embeddings, mSE, multilingual semantic embeddings |
| Liên quan | 5 | 5 |
| Tóm tắt≠ | Multilingual topic modeling extends probabilistic topic models such as LDA to corpora spanning two or more languages, inferring shared latent topics across language boundaries. By tying topic distributions across languages, it enables cross-lingual document analysis, comparable topic discovery, and information retrieval without requiring full parallel corpora. | Multilingual sentence embeddings map sentences from many languages into a single shared vector space so that semantically equivalent sentences — regardless of language — land close together. Models such as LaBSE, multilingual Sentence-BERT, and mUSE have made it practical to compare, retrieve, and classify text across 50 to 100+ languages without translating anything first. |
| ScholarGateBộ dữ liệu ↗ |
|
|