Tagged “data-quality”
-
Why the tax amount is harder to extract than the total
Every receipt shows a total. Not every receipt shows its tax as a number you can read. Six reasons the tax field fails, and what to do instead.
-
What a confidence score actually tells you
A confidence score is a model's estimate of its own reading, not the odds it is right. What it measures, and how to set a threshold on it.
-
Designing a review queue people actually use
Automated capture still needs human eyes, but only on the right items. What belongs in a review queue, and what should never reach one.
-
Duplicate receipts are the most common data problem
Most receipt data quality issues are not misread amounts. They are the same expense entered twice, and the same expense in two forms.
-
Vendor names are the hidden data quality problem
One supplier appears under five different names across your records. Nothing looks broken, and every vendor-level report is quietly wrong.
-
When you need line items and when the total is enough
Line-item extraction is the least reliable part of receipt scanning. Four situations justify the trouble, and most expenses are not among them.
-
A quarterly check on your own expense data
Individually correct entries can add up to wrong reports. Six aggregate checks that find the errors no per-receipt review will ever catch.