Groundeddocs

Parsing and OCR

Turn on OCR for scanned documents, choose a backend, test it, and deal with documents that failed or need OCR.

Grounded parses PDF, Word, PowerPoint, HTML, Markdown and text itself. OCR reads what has no text layer: scanned PDF pages, and PNG, JPEG and TIFF uploads. It's off by default. With it off, parsing is exactly as without OCR: a PDF with no text is skipped as "Needs OCR", and images can't be uploaded.

How it works

  • Only pages without a text layer are read. Grounded renders each such page to a greyscale image (300 DPI) and sends it to one backend. The text goes under the page's marker, so page numbers and citations work as for any page.
  • An image upload is a one-page document. Only the first page of a multi-page TIFF is read, with a warning.
  • A document's page says which pages were read, for example "Pages 3–7 were read with OCR (Tesseract)".
  • The crawler still ignores images.

Choosing a backend

BackendDeployGood forCost
Tesseract (recommended)The ocr-tesseract component (the grounded-ocr sidecar image)Printed text in the configured languages; fast on CPUsCPU only
Apache TikaThe tika component with Tika's -full imageInstalls that already run TikaCPU; the full image is large
Vision modelA model of kind Vision (OCR) under ModelsForms, tables and complex layouts; returns MarkdownTokens per page, recorded in the usage ledger

Only configured backends can be chosen: Tesseract needs OCR_TESSERACT_URL, Tika needs TIKA_URL, and the vision backend needs a vision model. Deploying the sidecar is covered in OCR sidecar.

A vision model is subject to its maximum classification, like any model: sources classified above it get no OCR (their scans stay "Needs OCR", and images can't be uploaded to them). Choose a model approved for the data you scan, or use Tesseract, which runs inside your cluster.

Turning it on

Choose the backend

Content → Parsing & OCR: turn on Read scanned pages and images with OCR, choose the backend and the Languages (Tesseract codes joined with +, such as eng or eng+spa), or the vision model. List only the languages your documents use: more languages are slower and can be less exact. A vision model reads any language.

Test it

Test reads a built-in sample page with the backend on the form, saved or not, and shows the text, the time and the confidence (for a vision model, its tokens). A failure shows the backend's error.

Save

Changes are audited. Each source also has its own OCR switch (on by default), so teams can turn it off where scans are noise.

Admin, Parsing & OCR: OCR turned on with the Tesseract backend and its languages.

Documents that failed or need OCR

The same page lists documents that failed or need OCR, grouped by team and source: the count, the reason (needs OCR, OCR error, damaged or unsupported file, other failure) and the oldest date. It never shows file names, titles or text. For each group, platform admins can:

  • Retry these: queue them again. Needs-OCR documents and OCR errors are refused, with the reason, while OCR is off for the platform or the source, or the vision model isn't approved for the source's classification. Damaged files usually fail again; the team has to fix and re-upload them.
  • Notify owners: the team's owners get Documents need attention, in the app and by email, with the count and a link to the source's documents filtered to Failed or Needs OCR. They can't turn this notification off. Shared sources have no owners to notify.

Both actions are audited. Auditors see the list but not the actions. Teams can also retry their own needs-OCR documents from the source's Documents tab.

Bounds

  • Per document: at most OCR_MAX_PAGES_PER_DOCUMENT pages (200); the rest are skipped with a warning.
  • Per worker: at most OCR_CONCURRENCY pages at once (2), across every document, so OCR can't take all of ingestion.
  • Per team per day: the team limit OCR pages per day (1,000, UTC day), under Limits. A document that would pass it waits and continues after midnight UTC, or at once when the limit is raised. 0 blocks OCR for the team.
  • Pages are rendered at 300 DPI, lower for very large pages, so no image is more than 6,000 pixels on its longest side. Image uploads over 100 megapixels are refused.

Every document read with OCR records its pages in the usage ledger, with the backend; a vision model also records its tokens, which costs can price. Tesseract and Tika pages are counted, not priced.

Troubleshooting

SymptomCheck
Documents fail with ocr_unavailableThe backend didn't answer (network, 5xx, timeout). They're retried, then fail; retry them once the backend is back. Check the sidecar's pods and its NetworkPolicy.
"OCR could not read page N"The backend refused that page; the rest of the document is indexed.
Scanned pages still skipped with OCR onThe source's OCR switch is off, the team's daily limit is 0, the vision model is disabled, or the source is classified above the vision model's level.
Poor textCheck the languages. Scans below about 200 DPI, or photos at an angle, read badly with Tesseract; a vision model does better on forms and tables.

On this page