Document Intelligence

Persian–English OCR

Certified-translation pipeline.

The problem it solved

Certified-translation offices drowned in manually retyping mixed Persian–English scans; quality and turnaround suffered. This pipeline extracts clean, structured text from low-quality, mixed-script documents automatically.

Similar problems it can solve

Any high-volume document-intelligence need: invoices and IDs, legal/medical record digitisation, KYC onboarding, and other RTL/LTR mixed-script extraction (Arabic, Hebrew, Urdu).

Multi-generation: PaddleOCR → DeepSeek-OCR → Qwen2-VL → Gemini 2.5 Flash post-processing. Page-aware chunking, SignalR progress, SQL Server jobs, C# OcrClient lib.

Overview

A production OCR service for Persian–English certified-translation documents. It ingests scanned immigration paperwork and returns clean, structured text, evolving through several model generations to handle mixed-script, low-quality scans.

Technology

A FastAPI backend with a multi-generation pipeline (PaddleOCR → DeepSeek-OCR → Qwen2-VL → Gemini 2.5 Flash post-processing), SQL Server job tracking, SignalR live progress, client authentication with quota/rate-limit, and a C# OcrClient consumption library.

FastAPI Gemini 2.5 SignalR