← Library · Frontier

Baidu's 'Unlimited OCR' Processes Dozens of Pages in a Single Pass

Baidu researchers have developed 'Unlimited OCR,' a new optical character recognition model that can process dozens of document pages in a single inference pass without increasing memory use or impacting speed based on text length. This is achieved through a redesigned attention mechanism called Reference Sliding Window Attention (R-SWA), which keeps the KV cache (a buffer for processed tokens) constant. The model, built on Deepseek OCR and utilizing a Mixture-of-Experts architecture, achieves 93.92% on the OmniDocBench v1.6 benchmark and significantly reduces error rates even on long documents.

Why it matters

This innovation overcomes a critical bottleneck in processing long documents, making OCR more efficient and scalable for applications involving extensive textual content. It offers a new approach to managing memory in large language models.

Learn one new AI thing every day.

Daily Deck sends you seven plain-English cards like this every morning. Free.

Start free