Skip to content
Kentron Technologies

AIAutomationDocuments

Reading documents with AI: invoices, prescriptions and the forms nobody wants to type

Document extraction is the automation with the clearest payback and the most overstated accuracy. What modern models do well, where they fail, and why the review step is the product.

Author
Kentron Technologies
Published
Reading time
6 min read
A mobile app home screen

Of all the things a language model can do for an Indian business, reading documents has the clearest payback. Purchase invoices arrive as PDFs and get typed into accounting software. Prescriptions arrive on paper. Forms arrive as photographs taken at an angle in bad light. Someone is typing all of it, and that someone has a better use of their time. This article covers what document extraction genuinely does well now, where it still fails, and how to build it so the failures are caught.

What changed, and what did not

Traditional OCR converted pixels to characters and left you with a wall of text, which meant the hard part, working out which number was the taxable value, was still a programming problem solved with brittle rules per vendor layout. Vision-capable models changed that: they read a document and return structured fields, so a layout they have never seen is handled without a template.

What has not changed is that documents are adversarially messy. A faded thermal print, a stamp over a total, a handwritten correction in the margin, two invoices photographed on one page. The model is dramatically better than rules at these, and still wrong often enough that the design question is not how to avoid errors but how to catch them.

Accuracy, stated honestly

DocumentRealistic expectationWhat breaks it
Clean digital PDF invoiceVery high on printed fieldsMulti-page tables, line items split across pages
Scanned or photographed invoiceGood, needs review on totalsSkew, shadows, thermal fading, stamps
Handwritten prescriptionPartial at bestHandwriting, abbreviations, drug name ambiguity
Structured formHighTick boxes, overflow writing outside the box

Be suspicious of any vendor quoting a single accuracy percentage. Accuracy depends on the document, the field and the photograph, and 'ninety-nine percent' usually means per character on clean input, which tells you nothing about whether the GST amount was right. The number that matters is per-document, per-critical-field, measured on your documents.

Design the review step first

The most common mistake is to build extraction and treat the human check as an afterthought. Invert it: the review interface is the product, and extraction is what makes it fast.

  • Show the extracted value next to the region of the document it came from, so a reviewer verifies by glancing rather than hunting.
  • Have the model report a confidence per field, and route low-confidence fields to review rather than the whole document.
  • Validate arithmetic. On an invoice, line items should sum to the subtotal and tax should be a legal rate. Most extraction errors fail an arithmetic check, which is free to run.
  • Never auto-post a financial document on first extraction. Review once, then relax the rule for layouts that have proven themselves.
  • Keep the original. Always. The document is the evidence; the extraction is a convenience.

Arithmetic validation is the single highest-value check available, because it catches the errors that matter most at no model cost. If the line items do not sum to the total, something was misread, and you know it before a human looks.

The model's job is to make review fast. It is not to remove review.

Where it pays for itself

The arithmetic from our automation method applies directly. A distributor receiving two hundred purchase invoices a month, each taking four minutes to enter, spends over thirteen hours a month on typing, plus the cost of the errors that reach the GST filing. Extraction with a review step typically turns four minutes into well under one, and the error cost drops because arithmetic validation catches what a tired person at six in the evening does not.

The GST context sharpens this. A mistyped taxable value does not just produce a wrong ledger; it produces a mismatch in a return, which is expensive to find and tedious to correct. Catching it at entry is worth more than the typing time saved.

The cost structure

Extraction is billed per document, as input tokens for the image or PDF plus output tokens for the structured fields. Larger and higher-resolution images cost more, which creates a genuine trade-off: downsample too aggressively and small print becomes unreadable, send full resolution and you pay for pixels that did not help. For a volume operation it is worth measuring where accuracy stops improving and setting resolution there.

A practical saving is to route by document type. A clean digital PDF often needs no vision model at all, since its text layer can be read directly and only interpreted. Reserve the expensive path for documents that are genuinely images.

Prescriptions, and knowing when to refuse

Handwritten prescriptions are the request we are most cautious about. A misread drug name or dose is a patient safety issue, not a data quality issue, and a model that is usually right is not an acceptable standard. Where we have worked with clinical documents, the design has been to assist a human rather than replace one: structure what is legible, flag what is not, and leave the authority with the person qualified to hold it.

That is the same principle behind how Healthixio handles dictation. A doctor's voice note becomes a draft discharge summary which the doctor then approves. The model removes the typing; the clinician keeps the decision.

Frequently asked questions

How accurate is AI invoice data extraction?

On clean digital invoices, very accurate on printed fields; on photographed or scanned documents, good but requiring review on totals; on handwriting, partial at best. Treat any single headline accuracy figure with suspicion, because it usually means per-character accuracy on clean input. Measure per-document accuracy on your own documents and on the specific fields that matter, such as the taxable value and the tax amount.

Can AI replace manual data entry completely?

Not for financial documents, and you should not want it to. The realistic and valuable goal is to make review fast rather than to eliminate it: the model proposes the values, the interface shows each one beside the part of the document it came from, and arithmetic validation catches the errors that matter. That typically cuts entry time by most of its original length while improving accuracy over manual typing.

What is the best way to validate extracted invoice data?

Arithmetic. Line items should sum to the subtotal, tax should be a legal GST rate applied correctly, and the total should reconcile. Most extraction errors fail one of these checks, they cost nothing to run, and they catch the mistakes with real financial consequences. Add per-field confidence from the model so only uncertain fields go to a human.

Is it safe to use AI to read prescriptions?

Only as an assistant to a qualified person. A misread drug name or dose is a patient safety matter rather than a data quality one, so the correct design structures what is legible, clearly flags what is not, and leaves the clinical decision with the clinician. We apply the same rule to dictation: the model drafts, the doctor approves.

Kentron Technologies

Editorial team

Builds and runs Kentron Technologies’s products. Writes here when a decision was hard enough to be worth explaining.

Next step

Tell us what you are running, and what is slow.

A demo of any product, or a conversation about something that does not exist yet. Either way, you will talk to someone who builds the software.

CallWhatsAppTalk to us