What is OCR document ingest?
OCR document ingest is the process of reading text out of scanned files and PDFs — including images of pages — and turning it into structured data you can query. iDBQuery can ingest a whole folder of documents this way, making their contents part of your one live model.
A lot of important information lives in documents, not databases — invoices, contracts, statements, scanned forms. OCR (optical character recognition) is the technology that reads the text out of those files, including scanned images where the text isn't selectable, so it can be analysed like any other data.
iDBQuery uses OCR ingest to make documents queryable. You can point it at a single PDF or an entire folder of files, and it processes them, extracts the content, and folds it into the same live model as your databases and spreadsheets.
With OCR document ingest you can:
- Ask questions of documents — query the contents of PDFs and scans in plain language, just like a table.
- Process folders in bulk — ingest many files at once rather than handling them one by one.
- Combine documents with structured data — a single question can join a figure from a PDF with records from a database.
The ingest runs in the background so large folders don't block you, and the extracted content is indexed so it's searchable and citable — meaning an answer drawn from a document still traces back to the file it came from. It's how unstructured paperwork becomes a first-class part of your analysis instead of something you read manually.
Updated 2026-06-22