Vertical notes
AI for accounting firms: a field report
We mapped a typical founder-led practice process by process, then graded AI against an answer key. Half the firm's week is work AI can mostly carry. The other half should stay exactly where it is.
Open the visual report — same data, built to skim.
You can't hire your way out of tax season. The US has roughly 340,000 fewer accountants and auditors than it did in 2019 — a ~17% decline — accounting degrees just posted their steepest one-year drop on record, and the AICPA estimates about three-quarters of CPAs are at or near retirement age. The way through isn't more headcount. It's taking the grunt work off the desks you already have.
Where the hours go
We started where every assessment should: the org chart. Who is paid to do what, hour by hour. Below is a representative model of a 16-person practice — in a real engagement, every row is measured, not modeled.
| Activity | Share of firm hours | What AI can carry |
|---|---|---|
| Return preparation | 24% | ~45% |
| Data entry into tax software | 17% | ~85% |
| Document intake & chasing | 14% | ~70% |
| Client communication & status | 12% | ~60% |
| Admin & scheduling | 10% | ~65% |
| Review & sign-off | 9% | ~15% |
| Advisory | 8% | ~10% |
| Business development | 6% | ~20% |
Representative 16-person firm. "What AI can carry" = the share of that activity's hours where graded performance clears the firm's bar. Estimates until a pilot measures them.
Half the firm's week sits in the top four rows — sorting documents, typing numbers in, chasing paperwork, answering "where's my return?" Advisory, the work clients actually pay a premium for, gets 8%. Nobody gets cut in this plan. The hours move up the table.
Don't trust AI. Grade it.
The rule is simple: AI does the firm's past work over again, gets graded against what the team actually did, and touches nothing until it passes.
So we ran the first grading pass — 40 cases on made-up tax documents (no client data), scored automatically on the Recursiv platform, July 29, 2026:
| The work | Right | Handed back | Cost each | Verdict |
|---|---|---|---|---|
| Sorting client documents 20 docs: W-2 · 1099 · K-1 · receipts · traps | 95.0% | 0% | $0.005/doc | On, with review |
| Typing the numbers in 20 boxes, graded one box at a time | 85.0% | 10% | $0.005/box | Keep practicing |
| Status replies & doc chasing | — | — | — | Measured in pilot |
| First draft of a 1040 | — | — | — | Measured in pilot |
| K-1 tie-out | — | — | — | Measured in pilot |
| Tax research memo | — | — | — | Measured in pilot |
Right = matches the answer key. Handed back = the AI gave it to a person instead of guessing. Bar: 98% to run alone, 90% to run with review.
Read the failures, because they're the product. Both hand-backs were smudged scans — and handing them back was the correct call. One document sort was right but garbled its reply format. And one field extraction fumbled a parenthesized loss, reading (12,340) the way a first-year associate might. Neither task cleared the 98% bar on day one. That is exactly why you grade before you trust — and why the verdict column, not a vendor's promise, decides what gets switched on.
Client data never leaves the building
Tax data lives under IRC §7216 and §6713, the FTC Safeguards Rule, IRS Pub 4557, and Circular 230. So the deployment is built so the default answer to "what leaves our environment?" is nothing: Recursiv runs on the firm's own servers, the AI comes to the data, and every action lands in an audit log mapped to the firm's WISP. Outside AI services are a consent-gated choice, off by default.
The path
- Walk the floor — who does what, hour by hour (~2 weeks).
- Practice runs — AI works alongside the team on real jobs; clients never see it (~4–6 weeks).
- The scorecard — what passed, what didn't, against the firm's own bar.
- Turn on what passed — behind review gates. Everything else keeps practicing until the numbers earn it.
If you run a firm — tax or otherwise — and want your version of this scorecard measured on your own work, that's what a pilot is. Get in touch.
Sources
- US Bureau of Labor Statistics, as reported by The Wall Street Journal (2023): ≈340,000 fewer accountants and auditors vs. 2019.
- AICPA, "Trends in the Supply of Accounting Graduates" (2023): bachelor's completions down 7.8% in 2021–22.
- AICPA estimate: ~75% of CPA members at or near retirement age.
- Scorecard rows 1–2: live graded run on the Recursiv platform, July 29 2026, 40 synthetic cases, $0.21 total inference cost. No client data.
Quick answers
How accurate is AI at accounting-firm paperwork today?
In our July 2026 graded run, AI sorted client tax documents at 95% and extracted form fields at 85%, at about half a cent per item. Neither cleared the 98% run-alone bar, so both run with human review.
Will AI replace accountants?
No. The target is the paper layer — document sorting, data entry, status emails — roughly half of a small firm's hours. Review, advisory, and IRS defense stay human.
How do you know when AI is safe to use on client work?
Grade it. AI redoes the firm's past work against an answer key in practice runs, and a task goes live only after it holds the firm's bar — for example 98% right with under 5% handed back.
Does client tax data leave the firm?
No. The system runs on the firm's own servers under IRC §7216, the FTC Safeguards Rule, and IRS Pub 4557. Outside AI services are off by default behind a consent gate.