view article Article LEMUR and Mean Centering for Late-Interaction Retrieval in txtai NeuML • 1 day ago • 5
Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders Paper • 2607.25180 • Published 18 days ago • 2
view article Article GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval chungimungi • 7 days ago • 11
Compact Language Models via Pruning and Knowledge Distillation Paper • 2407.14679 • Published Jul 19, 2024 • 43
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search Paper • 2607.27178 • Published 17 days ago • 6
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 10 days ago • 23
view article Article After the party comes the free lunch: regularizing ColBERT models to enhance pooling capabilities and reduce index footprint lightonai • Jul 6 • 15
view article Article mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models lightonai • 16 days ago • 36
jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation Paper • 2607.18152 • Published 26 days ago • 4
ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking Paper • 2506.03487 • Published Jun 4, 2025 • 7
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 30 days ago • 59
view article Article Beyond LoRA: Can you beat the most popular fine-tuning technique? +2 BenjaminB, sayakpaul, hubnemo, kashif • Jun 18 • 93
Training Sparse Mixture Of Experts Text Embedding Models Paper • 2502.07972 • Published Feb 11, 2025 • 12
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
view article Article Party is over: regularizing ColBERT models to fix efficient ANN methods lightonai • Jun 16 • 23
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking Paper • 2405.07920 • Published May 13, 2024 • 4
F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World Paper • 2603.19223 • Published Mar 19 • 39
Is Position Bias in Dense Retrievers Built In-or Learned from Data? Paper • 2605.26578 • Published May 26 • 21