How does iDBQuery read PDFs and scanned documents?
iDBQuery reads PDFs and scanned documents by extracting their text and structure — using OCR for scanned or image-based pages — and turning the content into part of your live queryable model. You can then ask questions across your documents the same way you query a database, with answers cited back to the source page.
Documents hold a huge amount of business data — contracts, invoices, statements, reports — that's normally locked away from analysis. iDBQuery unlocks it.
When you add a PDF or document, iDBQuery:
- Extracts the text from native PDFs directly.
- Runs OCR on scanned pages and image-based PDFs, reading the characters off the page so even a photographed or faxed document becomes searchable, structured data.
- Captures structure — tables, fields and values — so figures inside a document can be queried, not just keyword-searched.
- Folds it into your live model, alongside your databases and spreadsheets, so a single question can draw on documents and structured systems together.
Because every extracted value keeps a link to where it came from, answers drawn from documents are cited back to the source — you can see the exact page and record behind a number. This is what lets you ask things like "what's the total across these invoices" or "which contracts mention this clause" and get a real, verifiable answer instead of digging through files by hand.
Updated 2026-06-22