LightOn Expands Its OCR Model to Arabic to Target the Middle East
On Friday, LightOn showcased the ability of its document understanding model, LightOnOCR-2, to process Arabic through fine-tuning. This expansion follows the launch of Console, its document platform, and specifically targets organizations in the Middle East dealing with archives and administrative documents.
Demonstration on 12,000 Synthetic Arabic Pages
LightOn tested the adaptation of LightOnOCR-2 to Arabic using a dataset comprising 12,000 synthetic pages and their reference transcriptions. The corpus covers various documentary scenarios: scanning artifacts, font variations, resolution levels, and diverse document types.
The output retains the training format of the model, with bounding box detection linking text to its spatial position. This demonstration relies on an internal pipeline for generating synthetic data, designed to cover languages currently underrepresented in market OCR tools.
Specific Challenges of OCR in Arabic
Applying optical character recognition to Arabic presents distinct challenges. The script is written from right to left, characters are connected in cursive form, and open data sets as well as specialized models are less available than for Latin alphabet-based languages.
For organizations dealing with Arabic archives, administrative, legal, or heritage documents, these limitations slow down the automation of document processing workflows. LightOn indicates that this demonstration meets identified needs in the Middle East, where the company is already working with public and private organizations.
Accessibility and Open Source Model
LightOn provides the necessary guides to reproduce this fine-tuning on its Hugging Face space, making this approach accessible to a wider audience and adaptable to other documentary contexts.
LightOnOCR-2 is released as open source under the Apache 2.0 license and plays a central role in the document ingestion process within Console, the company's self-service offering. The base model achieves a score of 83.2% on the OlmOCR-Bench benchmark.