LightOn is a leading European generative AI company delivering secure on-prem RAG for document intelligence, enabling safe use of sensitive data behind firewall

Paris, France
LightOn × @infocom94 : Rencontres du numérique territorial. Pour cette première édition, @infocom94 a réuni plus d’une centaine de participants, parmi lesquels des acteurs technologiques et des représentants des collectivités, pour partager leurs problématiques, échanger sur leurs pratiques et mutualiser leurs connaissances et outils technologiques. En tant que partenaire d’Infocom 94, notre Head of Customer Services & Solutions, Gauthier Zuppinger, a pris la parole lors de la table ronde sur la souveraineté numérique pour aborder ce que signifie concrètement disposer d’une technologie souveraine, ainsi que les risques et les enjeux liés à l’utilisation de technologies non souveraines. L’ambition d’Infocom 94 d’apporter aux collectivités des solutions réalistes, pragmatiques et frugales rejoint pleinement nos valeurs et l’ADN de LightOn. Merci à Vincent Kerbiquet, à @AmbroiseToin président d’Infocom 94, et à TerikSene pour leur confiance et leur invitation à cette première édition !
1
3
302
Building AI search is one thing. Maintaining it in production is another. 🔍 New formats, permissions, retrieval quality, infrastructure, maintenance... and what happens when the people who built it leave? 🧠 Our latest article looks at the real cost of building enterprise AI search in-house, and when it makes more sense to buy. With LightOn Console, the search infrastructure is already there. Your team can focus on the application. 👉🏻 Read the article: lighton.ai/lighton-blogs/bui… 👉🏻 Explore Console: console.lighton.ai/
1
2
328
Une application IA, construite étape par étape et en direct. Le 29 septembre à 11h, @LightOnIO × @ScalingoHQ vous proposent une démonstration technique basée sur des documents médicaux fictifs mais réalistes. Au programme : 🔍 Extraire et structurer les données 🧠 Retrouver et contextualiser les informations utiles 📄 Générer automatiquement un dossier de sortie ☁ Déployer l’application et ses bases de données dans un environnement adapté aux données sensibles Une démonstration concrète pour comprendre comment assembler chaque brique et passer du cas d’usage à l’application déployée. 🎥 Démo live suivie d’une session de questions-réponses Inscription gratuite : app.livestorm.co/scalingo/de…
1
2
5
488
🤝 @eurodecision × @LightOnIO : allier science de la décision et IA générative. LightOn s’associe à @eurodecision , expert en Mathématiques Décisionnelles, Recherche Opérationnelle et IA, pour aider les organisations à mieux exploiter leurs données et leurs connaissances et à prendre de meilleures décisions grâce à l’IA. 🎯 L’objectif ? Combiner l’expertise d’EURODECISION en modélisation et optimisation avec celle de LightOn en IA générative pour développer des outils d’aide à la décision plus performants, plus rapides à déployer et plus simples à utiliser. Cette collaboration s’inscrit notamment dans la dynamique du PIIEC-IA, qui vise à accélérer l’intégration de l’IA dans l’industrie et à faire émerger une offre européenne d’IA à forte valeur ajoutée. Concrètement, plusieurs cas d’usage sont au cœur de cette collaboration : - Gestion des connaissances : analyser et exploiter des documents techniques, spécifications, normes, rapports de maintenance ou données métier. - Ingénierie logicielle : analyser les exigences et accompagner la rédaction de documentation technique et la génération de tests. - Modélisation & optimisation : identifier les modèles et approches adaptés à des problématiques de planification et d’optimisation sous contraintes. - Analyse & aide à la décision : expliquer les résultats des modèles en langage naturel et comparer différents scénarios « What-If ». Un partenariat pour transformer des problématiques complexes en décisions concrètes. 👉🏻 eurodecision.com/ 👉🏻 lighton.ai/ #IA #IAGenerative #AideALaDecision #RechercheOperationnelle #Optimisation
1
5
622
Your documents are not general-purpose, Your retriever should be good beyond these @tomaarsen recently fine-tuned mLateOn-unsupervised for medical retrieval using MIRIAD, a benchmark with 1,000 queries and 200,000 passages. Out of the box, mLateOn reached 0.8520 NDCG@10. After a short fine-tuning run on a single RTX 3090, mLateOn-medical reached 0.9139, outperforming every general-purpose retriever tested. Qwen3-Embedding-4B, the strongest dense model in the evaluation, reached only 0.7817, despite having roughly 33 times more active parameters. The difference comes down to how information is represented. Dense models compress an entire document into one vector. Late-interaction models preserve token-level representations, allowing them to capture signals that single-vector models may average away. This matters even more for long, specialized documents. In this evaluation, passages averaged 941 tokens. The conclusion is simple: strong retrieval is not only about building larger models. It is about choosing an architecture that can understand your documents and adapt to your domain. → Explore Tom Aarsen’s complete fine-tuning guide : huggingface.co/blog/train-mu… → Put advanced retrieval to work with LightOn Console console.lighton.ai
4
18
923
All your agent needs is search It gives them access to the right enterprise knowledge, grounds their responses in reliable evidence, and reduces hallucinations. But performance alone is not enough. For AI agents to operate at scale, they also need to make economic sense. An infinite context window is not an efficiency strategy. By retrieving only what matters, good retrieval reduces unnecessary context, lowers token usage, and makes every inference more cost-effective. Powerful agents are useful. Powerful and efficient agents are production-ready. 👉🏻 Start building with LightOn Console: console.lighton.ai
1
7
446
J-7 before the first Destination AI Coffee Chat! On August 27, LightOn joins @TDSYNNEX and @HPE for a practical conversation on turning AI ambitions into scalable enterprise solutions. 📅 August 27, 2026 🕦 13:00 to 14:00 Register here: connect.tdsynnex.be/event/de… #EnterpriseAI #SovereignAI #DestinationAI
2
675
RAG quality is decided before the prompt reaches the LLM. LightOn’s open-source stack strengthens every layer between raw documents and grounded answers: 📄 LightOnOCR-2 for parsing at scale 🔍 mLateOn, a state-of-the-art retriever 🎯 LightOn-rerank, a state-of-the-art reranker Better parsing. Sharper retrieval. Cleaner context. Use each component independently, or access the full production-ready pipeline through LightOn Console. The bricks are already in your dependency tree. Start building: console.lighton.ai/signup
1
5
17
982
Your AI doesn't need more context. It needs the right receipts, contract clauses, meeting transcripts and diagrams 📄 LightOn Search grounds every answer in your organization's truth by retrieving the most relevant evidence before it reaches your LLM. Less noise. Fewer tokens. Sourced answers from the documents that actually matter. Always. The best retrieval system doesn't find more. It finds what counts. ⚡ Start building with free credits: console.lighton.ai #RAG #Search #EnterpriseAI #Developers
1
1
4
472
Great search deserves a real test. Public datasets are now available directly in LightOn Console, so you can test our search on real-world document collections from the start. Explore the data, run semantic searches, ask questions, and assess retrieval performance without downloading, cleaning, chunking, or indexing documents first. Less setup. More experimentation. Faster prototypes. Start now 👉 console.lighton.ai/signup
2
7
567
From free tier to production in one click. Need more capacity? Upgrade your plan instantly, directly from the Console. No back-and-forth. No manual activation. Build, scale and deploy on your own timeline. 👉Get started: console.lighton.ai/signup
3
1
6
1,729
Drumroll please! After DenseOn and LateOn, we release mDenseOn & mLateOn: two fully open 307M retrievers for multilingual, long-context AND code search! Same backbone, same data, using translate-train. But very different behaviour once we leave the training languages 👀
Made with AI
1
13
41
1,439
New State-of-the-art Retrieval Model / Новая передовая поисковая модель / 全新最先进检索模型 / 최첨단 검색 모델 / 最先端の新しい検索モデル / نموذج استرجاع جديد متطور One retriever family for the languages your data actually speaks. Multilingual retrieval, built for production. A few months ago, LightOn open-sourced DenseOn and LateOn, its 149M-parameter retrievers. A new BEIR state of the art for their size class, Apache 2.0. Today, the multilingual pair: mDenseOn and mLateOn. 307M parameters. SOTA on BEIR, MLDR, MIRACL and MTEB Code. Nine languages, one quality bar: English, French, German, Spanish, Italian, Portuguese, Swedish, Norwegian, Arabic. Late interaction compounds. mLateOn stays competitive retrieving in languages and scripts it never saw in training: Cyrillic, Japanese, Korean, Chinese. It breaks the ceiling of translate-train pipelines. Everything open: models, datasets, training code. 👨‍🍳Kudos to @raphaelsrty @antoine_chaffin @paulomouraj @AmelieTabatta 🔗Read the full story: lighton.ai/lighton-blogs/mde… ⚡️Try it now: console.lighton.ai/signup
1
10
27
1,513
The LightOn Console SDK is here. One package to connect your applications and agents to state-of-the-art document parsing, retrieval, and extraction through a clean, Python-native interface. Install in seconds. Build in minutes. 📦 Get started → github.com/lightonai/lighton…
1
4
8
1,117
A few months ago @treygrainger and @softwaredoug invited me to guest lecture in the AI-Powered Search course! I have been a fan of the book for a long time, so I was beyond hyped and went deep in the multi-vector rabbit hole. The result is my 64-slide monster baby 😂
9
17
110
15,130
All you need is multi-vector! Better generalization. Better on long context. Better on long documents. That is late interaction: one representation per token instead of a whole document folded into a single point. Everything agents throw at retrieval, longer queries, longer inputs, domains no one fine-tuned for, is exactly what it was built for. @AmelieTabatta laid out the entire field in one lecture for @softwaredoug and @treygrainger's AI-Powered Search course. The closest thing to a field guide the space has. 🔗 Full lecture:  meet.ameliechatelain.com/lec…
2
15
51
5,081
Search designed to scale. Built for agents. LightOn Search reaches 86.27% accuracy with only 9.7 search calls on average, 15–25% fewer than the competitors ranked directly above us. Turns out brute force isn’t a retrieval strategy. Knowing exactly where to look is. Agent-era retrieval is here 🎯 Kudos to @raphaelsrty and @baptaubertin for running the benchmark and the whole LightOn Tech Team for making it happen 👉 Put LightOn Search to the test: console.lighton.ai 👉 See the benchmark on Hugging Face: huggingface.co/spaces/Tevatr…
5
11
908