Tencent Releases WeMM-Embedding Multimodal Models
Summary
Tencent has released WeMM-Embedding, a family of 2B, 4B, and 9B universal multimodal embedding models. The models support text, images, videos, visual documents, and interleaved inputs, and are trained through alignment and refinement stages. Tencent reports that the 2B variant beats a leading 8B open-source baseline on MMEB-v2, while the 9B model achieves an overall score of 80.6. The models also showed gains in internal evaluations and 14 online A/B tests, and are deployed across WeChat search, recommendation, content, and e-commerce applications. Model weights and code are available publicly.