What does it cost a language model to read a long document? Not in money — in memory. That question is the whole story behind a technical report Baidu posted this week, titled Unlimited OCR Works, and answering it properly tells you something about where document-AI systems break in production.
Start with the term. OCR — optical character recognition — is the task of turning an image of text into the text itself. For decades this was a pipeline of detectors and classifiers. The recent shift, exemplified by DeepSeek-OCR, is to treat OCR as a generation problem: a vision encoder turns the page into a compact set of vectors, and a large language model (LLM) decoder writes out the characters one token at a time, as the paper describes. The appeal is real. Because the decoder is a language model, it carries a prior over language — it knows that "recieve" is probably "receive", that a half-occluded word in a sentence has only a few plausible completions. That prior fixes errors a pure pixel classifier never could.
Now the cost. A transformer decoder generates each new token by attending to every token it has already produced. To avoid recomputing the past at each step, it stores intermediate vectors — the KV cache (key–value cache). The catch is structural: the cache grows linearly with the output length. Transcribe one page and the cache is small. Transcribe forty pages and the cache is enormous, memory consumption climbs, and generation slows down the further you get, exactly as the report notes. The model gets tired, in a sense — it spends more and more of its budget just remembering what it already wrote.
The authors frame this against a sharp observation: a human copying a long text shows no such slowdown. You do not re-read every previous page to write the next sentence. You hold a small working memory — the line you're on, a little context around it — and slide it forward. The decline that hits the transformer is not a law of nature; it's an artifact of how attention was wired.
Unlimited OCR's move is to rewire it. Taking DeepSeek-OCR as the baseline, the authors replace every attention layer in the decoder with what they call Reference Sliding Window Attention, or R-SWA, per the abstract. The idea is in the name. Instead of attending to the entire history, each step attends to a bounded window plus a fixed reference — which keeps the KV cache constant through the whole decode, rather than letting it grow. Combine that constant-memory decoder with the high compression rate of DeepSeek-OCR's encoder, and the model can transcribe dozens of pages in a single forward pass within a standard 32K maximum length.
That phrase — single forward pass — is the practical payoff. Anyone who has run document pipelines knows the usual workaround for long inputs: chunk the document, process each chunk, stitch the results, and pray the seams line up. Chunking is where errors hide. A table split across a page boundary, a footnote referenced three pages later, a heading whose scope spans the cut — these are precisely the cases that break naive chunking. A model that holds the whole document in one pass with flat memory cost sidesteps that entire class of bug.
I'd add a caution from building these systems. Constant KV cache is a memory guarantee, not an accuracy guarantee. Bounding the window is, by construction, throwing away long-range context — the model can no longer freely attend to something forty pages back. For copying tasks that's usually fine; the relevant context for transcribing a line is local. But for tasks that genuinely need long-range reference — resolving a pronoun against a name introduced much earlier, or reconciling a total against line items scattered through a report — a sliding window is a real constraint, and the "reference" component is presumably what's meant to soften it. How well it actually holds depends on numbers the abstract doesn't give. The report is a v1 with code and weights promised on Baidu's GitHub; the benchmarks are what I'd read before trusting it on anything dense.
The part worth lingering on is the authors' own claim that R-SWA is general. They note it applies beyond OCR — to automatic speech recognition (ASR), translation, and similar tasks, as stated in the abstract. That generality is not incidental; it points at the shared shape of these problems. OCR, speech-to-text, and translation are all parsing tasks: a long input stream maps to a long output stream, mostly in order, with strong local alignment between the two. When the alignment is local and monotonic, paying for global attention at every step is waste. You're carrying the whole transcript in working memory to write the next word, when the next word only needs the last few.
This is the functional lesson, and it generalizes past any single paper. The default transformer treats every token as potentially relevant to every other — maximally expressive, maximally expensive. Most real tasks don't need that. They have structure: locality, monotonicity, a natural notion of "what's nearby matters most". Architectures that bake that structure in — bounded windows, fixed references, constant caches — trade a slice of theoretical reach for resource behavior that doesn't degrade with scale. In production, behavior that doesn't degrade with scale is usually worth more than the slice you gave up.
The right way to read Unlimited OCR, then, is not as a better text scanner. It's a small, concrete instance of a recurring engineering choice: when your task has structure, stop paying for the generality you aren't using.
Sources
- quick links — Manual / ad-hoc · 2026-06-25