Documents / Data Extraction Live
OCR Data Extractor
Read a scanned invoice or receipt and pull out the header fields and line items as structured data.
- Run on demand
- ~5 credits per run
- v1.0.0
What does OCR Data Extractor do?
OCR Data Extractor is a production AI agent in the documents section of the Leverge agent store, built for the data extraction process. Read a scanned invoice or receipt and pull out the header fields and line items as structured data. It runs on demand, works through 4 steps and returns 4 outputs, including checks.
What it needs
-
Scanned document
A scanned or photographed PDF. Pages are rasterized and read with OCR.
-
Extraction settings
What it does
- Rasterizing pages
- Reading the scan
- Checking scan quality
- Checking the numbers add up
What you get back
- Checks
- Header fields
- Line items
- Scan quality
After each run it asks: “Did this read the document correctly?”
When it runs
Run on demand
Oversight
Runs under scoped, least-privilege credentials with every action written to an audit log. Anything that moves money, alters a contract or reaches a customer requires human approval before it executes.
Data Extraction
Other agents in data extraction
Extraction and classification across real-world file formats
-
Document Comparison
Compare two versions of a document side by side and list every substantive change, separating what alters meaning from what only alters wording.
-
Document Summarization
Upload a document and get a structured summary with key metadata and topics. Long documents are summarized section by section with rolling context.
-
Document Translation
Translate a document section by section, carrying terminology forward so the whole reads as one piece rather than a set of fragments.
Next Step
Deploy OCR Data Extractor, or adapt it
It runs as-is against the inputs above. Most deployments diverge — a different source system, a different tolerance, a different approval path. A 30-minute technical call establishes which.