TL;DR: BornoOCR (বর্ণ + OCR) is a browser-based PDF editor built in Bangladesh that solves the hardest problem in Bangla document technology: exact Bengali text recognition with format preservation. It uses a hybrid multi-engine OCR pipeline — PaddleOCR + Surya + Tesseract — with Bangla character-confusion correction, table detection, and Collabora Online editing.
The Bangla OCR Problem Nobody Solved
If you've ever tried to extract text from a scanned Bangla PDF, you know the pain. You upload a government form, a university transcript, or a legal document to Google Drive OCR — and what comes back is gibberish. শ and ষ get swapped. Conjunct consonants (যুক্তাক্ষর) like ক্ষ, ত্র, জ্ঞ break into fragments. Vowel signs (মাত্রা) detach from their consonants. The formatting is destroyed.
This is not a minor inconvenience. For a country of 170 million people where Bangla is the official language, the inability to reliably digitize Bangla documents is a systemic bottleneck. Government offices, universities, law firms, hospitals, and businesses across Bangladesh deal with paper and scanned PDFs daily — and have no reliable way to make them editable, searchable, or machine-readable.
The root cause is that Bangla (Bengali) is one of the hardest scripts for OCR engines to handle:
- Complex conjuncts: Two or three consonants merge into a single glyph (যুক্তাক্ষর). ক + ষ = ক্ষ. ত + র = ত্র. জ + ঞ + য = জ্ঞ্য. There are over 150 common conjuncts, and OCR engines must recognize each as a unit, not as separate characters.
- Context-dependent shaping: The same letter changes shape depending on its position in a word. র looks different at the start, middle, and end of a word. This confuses engines trained on Latin scripts.
- Vowel signs (মাত্রা): Bangla has vowel signs that attach above, below, before, and after consonants — ি (pre-base), ী (post-base), ু (below), ূ (below), ে (post), ৈ (post), ো (post), ৌ (post). OCR engines must correctly associate each matra with its base consonant.
- Character confusion: Visually similar characters are hard to distinguish: শ/ষ/স, র/ড, ন/ণ, ব/ব/ব, ড/ড়, ঢ/ঢ়. Tesseract alone confuses these constantly.
- Mixed scripts: Bangladeshi documents frequently mix Bangla and English — sometimes in the same line. An OCR engine must handle both scripts simultaneously.
The result? Standard OCR tools achieve only 60-75% character accuracy on Bangla. That means 1 in 4 characters is wrong. For a 1000-word document, that's 250 errors — unusable for any serious purpose.
How BornoOCR Solves It: A Hybrid Multi-Engine Pipeline
BornoOCR doesn't rely on a single OCR engine. Instead, it uses three engines with intelligent page-level routing, plus a dedicated Bangla post-correction layer. Here's how each engine works and when it's used:
1. PaddleOCR (PP-Structure) — For Tables and Forms
PaddleOCR, developed by Baidu, includes a layout analysis module called PP-Structure that excels at detecting tables, form fields, and structured layouts. For pages with tables, forms, or multi-column layouts, BornoOCR routes to PaddleOCR because:
- It detects table boundaries, row/column structure, and individual cells
- It recognizes form fields (checkboxes, text fields, key-value pairs)
- It handles Bangla text within table cells accurately
- It classifies page regions (text, title, list, table, figure, header, footer)
This is critical for Bangladeshi government forms, invoices, and academic transcripts — all of which are table-heavy.
2. Surya OCR — For Text-Heavy Pages
Surya OCR is a newer transformer-based OCR engine that excels at reading order, layout detection, and text-heavy pages. For pages that are mostly paragraphs of text (articles, reports, letters), BornoOCR routes to Surya because:
- It correctly identifies reading order in multi-column layouts
- It handles complex script shaping better than Tesseract
- It produces cleaner line and paragraph segmentation
- It works well on degraded or low-quality scans common in Bangladesh
3. Tesseract — Fallback and Client-Side
Tesseract is the oldest and most widely available OCR engine. While its Bangla accuracy is lower (60-75%), it runs in the browser via Tesseract.js — which means the Free tier of BornoOCR works entirely offline. Tesseract is used as:
- The client-side OCR engine for the Free tier (no server needed)
- A fallback when PaddleOCR and Surya are unavailable or fail
- A quick-preview mode before high-accuracy server OCR
4. Bangla Character-Confusion Corrector
This is BornoOCR's secret weapon. After any OCR engine produces raw text, a dedicated Bangla corrector post-processes the output:
- Confusion matrices: A pre-built matrix of commonly confused character pairs (শ↔ষ, র↔ড, ন↔ণ) with probability scores. The corrector checks each character against its context and swaps if the confusion probability is high.
- Dictionary matching: A Bangla word dictionary (50,000+ words) is used to validate and correct OCR output. If a word isn't in the dictionary but a 1-character-edit neighbor is, the corrector fixes it.
- Spacing correction: Bangla OCR often inserts or removes spaces incorrectly. The corrector fixes word boundaries using n-gram frequency analysis.
- Quality scoring: Each page gets an OCR confidence score (0-100). Pages below 70% are flagged for manual review in the OCR Correction Panel.
The result: 92-97% character accuracy on clean Bangla printed documents — a massive improvement over any single-engine approach.
Format Preservation: Not Just Text, But Layout
Extracting text is only half the problem. The other half is preserving the document's format — fonts, sizes, positions, tables, images, margins, headers, footers, and page numbers. Most OCR tools dump plain text. BornoOCR preserves the full document structure:
- Page layout: Multi-column layouts, margins, page sizes are preserved
- Tables: Detected tables become editable spreadsheet-like structures with cell-level styling
- Images: Embedded images, figures, and logos are kept in position
- Fonts: Bangla font references (Noto Bengali, Siyam Rupali, etc.) are preserved. Bundled OFL fonts ensure correct rendering.
- Headers/footers: Page headers, footers, and page numbers are detected and preserved
Collabora Online: Word-Style Editing in the Browser
BornoOCR integrates Collabora Online (CODE) — the same engine that powers LibreOffice Online — as its rich text editor. This means you get a full Word/Excel-like editing experience directly in your browser:
- Full text formatting (bold, italic, fonts, sizes, colors)
- Table creation and editing
- Page layout controls (margins, orientation, columns)
- Track changes and comments
- DOCX import/export with high fidelity
The workflow: Import a scanned PDF → OCR extracts text with format → export to DOCX → edit in Collabora → export back to PDF. The entire round-trip preserves Bangla text, fonts, and formatting.
Use Cases for BornoOCR in Bangladesh
Government and Public Sector
Bangladesh government offices deal with millions of paper forms — birth certificates, land records, tax forms, NID applications. BornoOCR can digitize these with high accuracy, preserving the form structure for database entry.
Academic and Research
Universities have decades of Bangla academic papers, theses, and journals in print. BornoOCR can digitize entire archives, making them searchable and citable. The literature discovery tool at Banglai Cloud Labs can then build citation networks from the digitized corpus.
Legal and Court Documents
Law firms and courts handle vast volumes of Bangla legal documents — contracts, judgments, petitions. BornoOCR's format preservation ensures legal formatting (numbered clauses, margin notes, stamps) is maintained.
Medical Records
Hospitals and clinics can digitize patient records, prescriptions, and lab reports. Combined with Banglai Cloud's Prescription Builder, this creates a complete medical document workflow.
Business and Finance
Invoices, receipts, bank statements, and financial reports can be digitized with table detection intact. The key-value extraction feature pulls structured data (invoice number, date, amount, supplier) automatically.
Newspapers and Publishing
Bangla newspapers and publishers can digitize archives going back decades. BornoOCR's multi-column layout detection handles newspaper layouts correctly.
BornoOCR vs. Alternatives: A Comparison
BornoOCR vs. Google Drive OCR
Google Drive OCR uses Tesseract under the hood with no Bangla post-correction. Accuracy: 60-75%. No format preservation. No table detection. No editing. BornoOCR: 92-97% accuracy, full format preservation, table detection, Collabora editing.
BornoOCR vs. Adobe Acrobat OCR
Adobe Acrobat has decent OCR but poor Bangla support. It's expensive ($20+/month), desktop-only, and doesn't include a Bangla character corrector. BornoOCR is browser-based, has a dedicated Bangla corrector, and costs ৳500/month — a fraction of Adobe's price.
BornoOCR vs. Tesseract (standalone)
Tesseract is free and open-source but achieves only 60-75% on Bangla. No format preservation, no table detection, no editing UI. BornoOCR wraps Tesseract as one of three engines and adds the correction layer that makes the output usable.
BornoOCR vs.ABBYY FineReader
ABBYY is a premium OCR tool ($200+) with decent Bangla support, but it's desktop-only, expensive, and doesn't include a browser-based editor or Collabora integration. BornoOCR is browser-based, cheaper, and purpose-built for Bangla.
Pricing: Accessible for Bangladesh
BornoOCR is designed to be accessible for the Bangladeshi market:
- Free tier: Basic PDF viewing, text editing, client-side OCR (Tesseract.js in browser), Bangla + English support, local project save (IndexedDB), PDF export. No login required beyond Banglai Cloud account.
- Pro tier: ৳500/month or ৳5000/year. Adds server-side hybrid OCR (PaddleOCR + Surya), Bangla OCR post-correction, table detection and editing, Collabora Online editor, high-fidelity PDF↔DOCX export, key-value extraction, barcode detection, and cloud project storage.
- 14-day Pro free trial: Every new user gets 14 days of Pro for free. No credit card required. After the trial, you're downgraded to Free — not locked out.
Payment is via bKash (Send Money), verified manually by our admin team — the same system used by TurfOS, Restaurant POS, and IMS on Banglai Cloud. This keeps costs low and accessible for Bangladesh.
How to Get Started
- Go to banglai.cloud/homemade-software/bornoocr
- Sign in with your Banglai Cloud account (free to create)
- Click 'Launch Editor' — you'll get a 14-day Pro trial automatically
- Upload a scanned PDF or image
- BornoOCR runs OCR automatically and presents an editable document
- Edit in the Collabora Online editor (Pro) or basic editor (Free)
- Export to PDF, DOCX, or save as a project
The Technology Behind BornoOCR
BornoOCR is built on a modern, production-grade architecture:
- Frontend: React 18 + TypeScript + Vite, Tailwind CSS, Zustand state, Fabric.js canvas, PDF.js rendering, Tesseract.js client-side OCR, HarfBuzz.js for Bangla font shaping, Dexie/IndexedDB for local persistence
- Backend: Fastify + TypeScript, PostgreSQL 16 with Drizzle ORM, Redis 7 with BullMQ job queue, JWT auth, multipart uploads
- Python OCR Worker: FastAPI, PaddleOCR + PaddlePaddle, Surya OCR, Tesseract, OpenCV, pdfplumber, Camelot, pyzbar (barcode detection)
- LibreOffice Worker: Unoserver for high-fidelity PDF↔DOCX/ODT/HTML conversion
- Collabora Online (CODE): WOPI-protocol integration for browser-based rich text editing
- Infrastructure: Docker Compose, Nginx reverse proxy, S3-compatible storage
The entire stack is designed to run on a single VPS (4GB+ RAM recommended) — making it deployable anywhere in Bangladesh.
Why This Matters for Bangladesh
Bangladesh is digitizing rapidly. The government's Digital Bangladesh vision, the rise of fintech (bKash, Nagad), and the growth of the IT sector all depend on one thing: the ability to turn paper documents into digital data. Until now, Bangla OCR has been the weak link.
BornoOCR closes that gap. It's not a foreign tool adapted for Bangla — it's built in Bangladesh, for Bangla, with the specific challenges of Bengali script in mind. From the character-confusion corrector that knows শ vs ষ vs স, to the table detector that handles Bangla government forms, to the Collabora editor that renders Noto Bengali correctly — every piece is designed for the Bangladeshi context.
This is what 'built from Bangladesh for the world' means. We start with the hardest problem — our own language — and build something that works. Then we share it with the 300+ million Bangla speakers worldwide.
What's Next for BornoOCR
BornoOCR is actively developed. The roadmap includes:
- Handwriting recognition for Bangla (currently print-only)
- Mobile app (Android) for document scanning on the go
- API access for developers who want to integrate Bangla OCR into their own apps
- Batch processing for large document archives
- More local language support (Chittagonian, Sylheti, Chakma scripts)
- On-premise deployment for government and enterprise
BornoOCR is available now at banglai.cloud/homemade-software/bornoocr. Try it free for 14 days — no credit card, no commitment. Just upload a Bangla PDF and see the difference.