Wire
UEmbed unifies sparse and dense retrieval
UEmbed-9B scored 71.8 for dense retrieval and 71.0 for sparse retrieval on MMEB-v2, producing both representations in one causal forward pass. The August 3 preprint releases 2B, 4B, and 9B decoder-only variants trained on public data, extending learned sparse retrieval across text and multimodal inputs without a separate cross-modal module. For builders weighing retrieval stacks alongside open-model price and deployment trade-offs, one shared model could simplify hybrid search, but the paper’s benchmark result still needs workload-specific checks on latency, index size, and relevance.