Accounting & Finance

How Techparser Built an AI Reviewer for Financial-Statement Packages

Techparser built an AI financial-statement reviewer for a CPA firm: deterministic arithmetic plus model review, 90 seconds and $0.24–0.28 per review.

Client
A CPA firm (name withheld)
Industry
Accounting & Finance
Timeline
Delivered
Team
1 senior full-stack engineer

Want results like these?




Results

$0.24–0.28
Measured model cost per review
~90 s
From draft package to exception report
~40 rules
Formatting rules read from the firm's style guide

The problem

A reviewer at the CPA firm checks every outgoing financial-statement package by hand: the arithmetic, the tie-outs between statements, missing disclosures and the firm's house formatting. Techparser built a tool that takes the draft Word report and its Excel workpapers and returns a categorised exception report and a ready-to-send PDF. Formatting rules could not be hard-coded; they had to come from the firm's own style guide, so the tool carries no conventions of its own. Uploads are large, so each request had to be bounded before anything was parsed, and the documents under review could never be stored. It also had to run the same on a laptop and in a container.

The stakes are a reviewer's signature. A model that states a wrong number confidently is worse than no check, because a reviewer who trusts it could sign off a defective statement. Enforcing a rule the firm never wrote creates busywork and erodes trust. And one failed model call must not sink the whole review.

What Techparser built

  • Code refoots every total, balances the balance sheet, rolls equity and cash forward, and ties ending cash and net income across the statements, within a one-unit tolerance.
  • The verified figures are passed to the model, and any model finding that contradicts them is dropped; the model judges disclosures and wording, not the numbers.
  • Each model finding must include a short verbatim quote, matched back to the document to give it a section and page number, so a reviewer can jump straight to it.
  • About forty formatting rules are extracted from the firm's uploaded style guide into a fixed schema; a rule the guide does not state stays empty and its check is skipped.
  • A new style guide is read and its rules saved before the old file is replaced, so a failed extraction leaves the previous guide working.
  • The model gets at most two attempts, and only for empty replies; if the review cannot run, the arithmetic, formatting and date findings still ship with a warning naming the cause.
  • Anything over about 51 MB is refused before the upload is spooled to disk, each review runs in a temporary directory that is deleted afterwards, and responses are never cached.
  • Sign-in tokens are verified locally against the provider's published keys, each request re-reads the user's role and blocked status, and row-level security limits what each user can read.
  • Input and output tokens are logged on every model call, which is how the $0.24–0.28 per review figure was measured on sample documents rather than estimated.

Decisions that mattered

Numbers belong to code, not the model

The core decision was to never let the model do arithmetic. A deterministic engine refoots every total, balances the balance sheet, rolls equity and cash forward, and ties ending cash and net income across the statements, within a one-unit tolerance. Those verified facts are handed to the model, and any model finding that contradicts them is dropped and noted. The model's job is judgement on disclosures and wording. The trade-off is more code to maintain than a single prompt, but a tool that confidently states a wrong number is worse than no tool, because a reviewer could sign off on it.

Rules from the firm's own guide

The firm's formatting is its own, so the tool carries no conventions of its own. About forty rules for fonts, sizes, capitalisation, currency signs and contents pages are extracted from the firm's uploaded style guide into a fixed schema. A rule the guide does not state stays empty, and its check is skipped rather than guessed. When a new guide is uploaded, its rules are read and saved before the old file is replaced, so a failed extraction leaves the previous guide working. Enforcing a rule the firm never wrote would create busywork and erode trust, so the tool only checks what the guide actually says.

Failure is designed, not hoped for

A review has to survive a bad model call. The model gets at most two attempts, and only for empty or unreadable replies; refusals and rate limits surface at once. If the review cannot run, the arithmetic, formatting and date findings still ship, with a warning naming the cause, so the reviewer is never left with nothing. Requests are bounded before parsing: anything over about 51 MB is refused before the upload is spooled to disk. Each review runs in a temporary directory that is deleted afterwards, and the documents under review are never stored. The trade-off is stricter limits up front in exchange for a tool that fails safely.

Outcome

The tool was delivered to the CPA firm. A draft financial-statement package goes in, and a categorised exception report and a ready-to-send PDF come back in about 90 seconds, at a measured $0.24–0.28 per review. It runs the same on a reviewer's laptop and in a container.

The design is what makes the tool trustworthy. Arithmetic is verified by code, so the model cannot confidently state a wrong number. Formatting is checked only against the firm's own style guide, so it never invents busywork. And when a model call fails, the arithmetic and formatting findings still ship, so a review is never lost to one bad request. Because nothing under review is stored, the firm's client documents leave no trace behind.

Questions about this project

What technology stack does the AI audit reviewer use?
The backend is Python and FastAPI, with Anthropic's Claude for the review step. Documents are parsed with python-docx, openpyxl and pdfplumber, and the PDF report is produced with reportlab, using LibreOffice for conversion. The frontend is Next.js and TypeScript, with Supabase for auth and data. It ships in Docker and runs on Railway and Vercel.
How does the tool avoid signing off on a wrong number?
Arithmetic is never left to the model. A deterministic engine refoots totals, balances the balance sheet and ties cash and net income across the statements within a one-unit tolerance. Those verified figures are passed to the model, and any model finding that contradicts them is dropped and flagged in the report. The model only judges disclosures and wording, so a confident but wrong number cannot reach the reviewer.
Can Techparser build an AI document-review tool like this for us?
Yes. Techparser built this reviewer end to end: the deterministic tie-out engine, the style-guide rule extraction, the model review with page-anchored findings, the PDF output and the auth and hosting. The same approach fits any document review where numbers must be exact and rules come from your own standards. Book a call to discuss your workflow.