Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

LightOnOCR 3 Adds Layout and Chart Extraction to Open OCR

Listen to this article

LightOn released LightOnOCR 3 on October 8, adding document layout, image descriptions and chart data extraction to its open OCR model family. The release announcement describes three variants under Apache 2.0, with a choice between plain transcription and output that preserves the location of content on the page.

A Bookeye scanner surrounded by books in the Cybernetics Library
A book scanner in the Cybernetics Library illustrates the document digitization workflow. Photo by ZhaoFJx under CC BY 4.0. Cropped by ByteForward.

For teams turning reports and scanned records into searchable knowledge, the useful change is that text can travel with its source location. That gives downstream software a way to connect extracted passages to the page regions a reader can inspect. It still leaves document validation, indexing and retrieval to the application.

Choose the model around your workflow

  • LightOnOCR 3 0.8B is the smallest option. It uses the Qwen3.5 vision language architecture and retains the family’s layout and visual extraction capabilities. LightOn positions it for efficiency.
  • LightOnOCR 3 1B keeps the LightOnOCR 2 architecture. That makes it the compatibility focused option for existing deployments. Its Transformers example uses the LightOnOcr classes and includes MPS, CUDA and CPU device selection. Those examples do not establish acceptable speed on every machine.
  • LightOnOCR 3 4B uses Qwen3.5 and is LightOn’s recommended variant for most OCR tasks. Treat the model name as its identifier rather than a promise about minimum memory or deployment cost.

Use the two supported prompt modes

Send a page image with no text prompt for transcription. Add only the word grounding for labeled content blocks and bounding boxes. The 4B model card explicitly warns that other instructions are outside the training distribution. Adding a long request for a custom schema is therefore a poor default integration strategy.

Grounded output includes brief image descriptions and chart values represented as HTML tables. Coordinates use a normalized scale from 0 to 1000. The model card reports about 25 percent more output tokens for grounding than plain transcription on its 4B sample. Budget for that extra generated output when comparing the modes.

What enters and leaves the pipeline

The official client and command line tools accept PDFs and images. Each processed page produces a PNG of the input and Markdown output. Grounding also produces parsed JSON blocks. Plain output can preserve HTML tables, LaTeX mathematics and reading order. A local viewer lets users inspect the results.

Preprocessing differs by variant. The 0.8B instructions render PDF pages at 400 DPI and resize to a longest edge of 2048 pixels. The 4B instructions use the same settings and disable thinking. Preserve aspect ratio. The 1B card instead recommends 200 DPI and a longest dimension of 1540 pixels.

Pin the serving environment

The repository’s vLLM setup pins vLLM 0.30.0 and Transformers 5.16.1. It warns that Transformers 5.17 breaks the 1B models with that vLLM version, while vLLM 0.27.1 causes looping. Follow the documented dependency combination before experimenting with upgrades. The client submits pages concurrently and lets vLLM batch requests, so throughput also depends on workload and serving configuration.

Read the benchmark claims with their context

LightOn reports 86.3 on olmOCR Bench and 75.1 on ParseBench for the 4B variant. Its release comparison shows different models leading different document categories. These are vendor measurements. ByteForward has not run independent performance tests.

There is also a material leaderboard caveat. The public ParseBench leaderboard lists KDL Frontier Parser nano at 76.36, above LightOn’s reported 75.1. That prevents an unqualified claim that LightOnOCR 3 is the overall leader.

LightOn’s reproduction repository records KDL at 72.39 locally and separately notes its published 76.36. It also lists the 4B model at 86.1 on olmOCR Bench with a postprocessed pipeline. These are different reported runs, not interchangeable figures. The repository documents pinned models, datasets and scorers so readers can investigate the conditions.

Evaluate the documents you actually process

For a first evaluation, use a small set of representative pages and keep the source image beside every result. The following checks matter more than selecting a model from its aggregate score alone.

  • Include degraded scans, dense tables, charts and multiple column layouts from the real workload.
  • Compare transcription and grounding for completeness, latency and output length on the same pages.
  • Manually verify numerical values, units and chart labels before feeding extracted data into decisions.
  • Check that each extracted block points to the correct source region and preserves reading order.

LightOnOCR 3 offers a practical reason to revisit document ingestion when source locations and visual content matter. The deployment decision should follow those checks, with enough review capacity to catch extraction errors before they enter a knowledge base.

For related coverage, explore ByteForward’s AI tools reporting.

A book scanner in the Cybernetics Library illustrates the document digitization workflow. Photo by ZhaoFJx under CC BY 4.0. Cropped by ByteForward.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile