← Library · Frontier

Baidu's 'Unlimited OCR' Processes Dozens of Pages with Constant Memory

Baidu researchers have developed 'Unlimited OCR,' an optical character recognition model that can process dozens of document pages in a single inference pass. It achieves this by using Reference Sliding Window Attention (R-SWA), which caps the KV cache, the memory buffer for past tokens, at a fixed size. This allows memory use and speed to remain constant regardless of the document's text length.

Why it matters

This innovation eliminates a major bottleneck in processing long documents, offering significant improvements in efficiency and cost for applications involving extensive text extraction, and could expand language model memory for long contexts.

Learn one new AI thing every day.

Daily Deck sends you seven plain-English cards like this every morning. Free.

Start free