LightOn launches LightOnOCR-3, a documentary AI that tops three benchmarks
In a press release issued on October 8, 2026 as regulatory information, the Paris-based publisher listed on Euronext Growth presented a family of AI models that goes beyond text recognition: image description, data extraction from charts and preservation of document structure. LightOn highlights a cost of use that it presents as "a fraction of the cost of competing models".
First on three open tests, second on OlmOCR-Bench
According to LightOn, LightOnOCR-3 ranks first on a set of three open reference tests evaluating text recognition and interpretation of visual information, ahead of existing OCR models in Europe, the United States and China. On OlmOCR-Bench, which the company describes as the industry's benchmark test for document text extraction, the model ranks second, behind Infinity-Parser2-Pro (a model more than eight times larger), and ahead of Chandra-OCR-2, Mistral OCR 4.1 and dots.mocr. LightOnOCR-3 also leads open-weight models on ParseBench, which evaluates layout rendering, image description and extraction of structured data from charts and scientific diagrams. It ranks first on FRBench-pdf2md, dedicated to French-language documents, with the largest gains in handwriting recognition. Detailed scores are referred to an associated technical article. The company targets three use cases: financial analysis (data extraction from text, tables and charts in reports), administrative processes (conversion of forms and correspondence) and search in technical documents, through identification of sections and information location on the page.
Two versions under Apache 2.0 license, up to twice as fast
LightOnOCR-3 succeeds LightOnOCR-2, which records 240,000 downloads per month on Hugging Face according to the company. The new family is offered in two versions of 0.8 billion and 4 billion parameters, to allow organizations to choose the model suited to their computing resources. According to LightOn, these models can process documents up to twice as fast as LightOnOCR-2 and present the lowest usage costs among major OCR models. Deployed on an organization's own infrastructure, these costs can fall below one cent per thousand pages, depending on hardware, usage level and workload. The combination of transcription, layout analysis and image processing in a single model aims to reduce the need to assemble separate tools. Distributed under Apache 2.0 license, the models can be used for commercial purposes, modified, redistributed and integrated into user applications.