BillConverter

A MeterID resource

Guides 5

Commodities Electric · Gas · Water · Sewer · Steam

AI Utility Bill Extraction Accuracy: How It Fails, and How to Catch It

There is no universal accuracy rate for AI bill extraction. Learn how to test your document mix and catch wrong values, missing meters, dates, and units.

There is no meaningful universal accuracy percentage for AI utility bill extraction. Results depend on scan quality, utility layouts, page count, requested fields, and whether accuracy is measured per field, per bill, or per meter row. A system that gets 99 of 100 requested cells right can still fail the bill if the missing cell is a meter row or total due.

Test on your own document mix. Build a labeled set that includes scans, multi-meter bills, estimated reads, credits, and long rate tables. Report exact-match accuracy by field, then separately report complete bills and complete meter rows. The second measure is usually more useful for operations.

Why valid-looking errors are hard to catch

Optical character recognition (OCR) often exposes character errors such as 1,2SO kWh, which a numeric type check can reject. A model may instead return 1,250 when the page says 1,350. The wrong value still has the right type and a plausible magnitude.

This can happen on a faint scan, beside a broken box border, or in a column of numbers under one shared header. Format and range checks pass, so the validation layer also needs arithmetic, document structure, and cross-bill history.

The failure modes

Confidently wrong values

Suppose a rescanned bill shows $3,418.28 inside a shaded amount-due box. A model returns $3,417.28. The result has valid currency formatting and a plausible magnitude. You catch it only by checking against another relationship on the page or reviewing the source.

This is an illustrative error, not a claim about a particular utility or model.

Right value, wrong field

Bills carry several money fields: current charges, previous balance, payments received, adjustments, and total amount due. When prior balance is zero, some of those values may match, so a field-mapping error can survive a small test set.

When a bill carries a prior balance and partial payment, current charges and total due diverge. Current charges belong in period-cost analysis; total due belongs in payment processing. Both values can be transcribed correctly and mapped to the wrong column. Define the fields before extraction; the field-by-field schema separates them.

Multi-meter bills flattened into one row

An account or summary bill can cover several services. PG&E uses Service Agreement IDs within an account. Georgia Power Summary Billing is available to qualifying commercial and industrial customers with at least ten eligible Southern Company accounts. National Grid’s Massachusetts tariff also describes summary billing for qualified customers with multiple electric accounts.

A model may read the summary and return one correct aggregate row instead of the service-level detail. Total due and usage can reconcile while meter-level history disappears. Compare extracted row count with the service or meter count on the bill. The identifier guide explains the account, meter, and service-level distinction.

Date confusion

Watch for three variants:

Statement date captured as service end. A bill might show a Jan 15 to Feb 13 billing period and a Feb 14 statement date. The statement date does not belong in service_end_date.

Overlaps and gaps between consecutive bills. A wrong boundary date changes both billing-day counts and can create an artificial swing in daily usage.

Periods spanning a month boundary. Assigning the whole entry to the statement month shifts usage away from the dates when it occurred. Preserve start and end dates, then calendarize only when the downstream report requires it.

Unit assumptions

Gas may be billed in therms, CCF, MCF, or another unit. Water may use CCF, kGal, gallons, or cubic meters. Do not infer a unit from the utility name alone; store the printed unit or mark it unknown.

The scale of the error varies. CCF read as MCF is 10 times too large or small. One CCF of water is about 748 gallons; one kGal is 1,000 gallons. Converting gas volume to energy requires a heat-content factor, so CCF and therms are not interchangeable.

Electric has its own version: demand in kW landing in the usage field, or kVAh captured as kWh on a bill that shows both.

Estimated reads recorded as actual

Utilities label reads as actual or estimated. National Grid’s bill guide, for example, explains both labels. Include the flag in the requested schema.

If estimates run low, the next actual read may absorb the difference. The resulting usage spike reflects billing correction as well as current-period consumption. Capture read type and review the sequence before treating it as a site anomaly.

Summary read instead of detail read

Some bills place the summary near the front and rate detail several pages later. If only the summary pages are processed, totals may be correct while line items and the supply-delivery split are absent. Record page count and require evidence for each requested section.

Hallucinated tariff codes

Rate schedules have recognizable formats, and utilities use different conventions. PG&E: A-1, A-10, B-19, B-10. Con Edison: SC 2, SC 9 Rate I, with a space rather than a hyphen. Georgia Power: TOU-GSD-17. SCE: TOU-GS-3.

An unclear code can become a valid-looking alternative: SC 9 Rate II instead of Rate I, or SC-2 instead of Con Edison’s printed SC 2. Check the value against the utility’s published schedules and the account’s prior bills.

Illustrative example: correct fields, wrong grain

The account, service agreements, meter numbers, and charges below are fictional. The example shows how a correct aggregate can still lose meter detail.

Service AgreementMeterkWhCharges$/kWh
5551234567101020304018,240$3,102.88$0.1701
555123456810102030416,410$1,204.55$0.1879
555123456910102030422,150$498.21$0.2317
Total26,800$4,805.64$0.1793

Current charges $4,805.64. Previous balance $1,918.42. Payments received $1,300.00. Total amount due $5,424.06.

Extraction returns: account 1234567890-1, service period 2025-02-06 to 2025-03-07, usage 26800 kWh, current charges 4805.64, total due 5424.06. Each returned field matches the illustrative bill. Arithmetic reconciles: 4,805.64 + 1,918.42 − 1,300.00 = 5,424.06.

The aggregate values are correct, but the grain is wrong. There is one row where there should be three, so the higher-cost common-area meter disappears into the blended rate. If one service later closes, there is no service-level history to explain the change. The aggregate row is also insufficient for a meter-level Portfolio Manager setup unless the user intentionally created an aggregate meter there.

A field-level score does not capture the missing rows. Measure row completeness separately.

The validation layer

Checks worth implementing, roughly in order of how much they catch per unit of effort.

Arithmetic reconciliation

Two identities, both cheap:

sum(line_items)                                        == current_charges   ± $0.02
prior_balance − payments + adjustments + current_charges == total_due       ± $0.02

Set a documented tolerance for rounding and compare it with real bills from each utility. Do not assume every mismatch above a fixed number is an extraction error; taxes and bill-level adjustments may need separate fields.

Cross-bill continuity

Compare each start date with the previous end date. Some bill sequences reuse the boundary date; others begin on the following day. Establish the convention per source and flag deviations. Review the first failures before automating rejection, because rebills and meter changes can create legitimate exceptions.

Usage sanity

Compare usage with the same period in prior years and with recent history. Choose thresholds from the account’s variability rather than applying one percentage to every site. Also flag changes near known conversion factors such as 10, 100, 748, or 1,000; those deserve a unit and decimal check.

Rate reasonableness

Derive cost per unit and compare it with the account’s own history or applicable tariff. This can expose an error in either usage or cost. Use a wider review band after a known rate change rather than relying on a national average.

Completeness

Maintain an expected bill schedule by meter and compare it with what arrived. A missing bill does not create a malformed row, so field-level checks cannot find it.

Duplicate detection

Use a business key such as account, meter or service ID, and service period in addition to the file hash. Two downloads of the same bill can have different file bytes. When the business key matches but totals differ, hold both records for rebill review instead of discarding one automatically.

Confidence and human review

Do not route review on model confidence alone. Review failed checks, changes to normally stable fields, new accounts or meters, and material dollar exposure. Measure the resulting review rate. If nearly every bill needs inspection, the workflow is not saving much labor and may need a narrower extraction scope or a different method.

Test the checks on labeled historical bills before using them to accept new data.

Doing this every month?

MeterID collects utility bills, checks the extracted data, and delivers structured output for multi-site portfolios — see how it works.