The short version
A typical founder-led firm, mapped process by process. AI does the work in the background, graded against what your team actually did. Only what passes gets switched on.
Includes a live graded run — real grades, made-up documents
52%
Half the firm's week is sorting documents, typing numbers in, and status emails. We graded AI on that work — 85–95% right on its first run (Fig. 3).
The whole report in three lines
Half your firm's week is paperwork. AI can carry most of it. The judgment work stays human.
We graded it: 85–95% right, half a cent per item. Real run, receipts below.
Nothing turns on until it beats your bar. Below the bar, it keeps practicing.
01 · Why now
Fewer accountants every year. Fewer coming. The way through isn't more headcount — it's taking the grunt work off the desks you already have.
Accountants who left the field
340K
Lost 2019→2023. The labor manual prep runs on is gone.1
CPA pipeline, one year
7.8%↓
Accounting degrees, one year. Steepest drop in decades.2
CPAs at or near retirement
~75%
Capacity you can't hire. You have to build it.3
02 · Your week
Representative 16-person practice. In the engagement, every row is measured, not modeled.
Fig. 2 — Where the hours go · % of all staff hours, representative model
Half the firm's hours sit where AI carries 60–85%. Advisory gets 8%. Nobody gets cut — the hours move to advisory, IRS defense, and winning new clients.
03 · How we test it
AI redoes your firm's past work. We grade it against what your team actually did. It touches nothing until it passes.
The answer key
200–500 pieces of your past work per task. What your team did is the right answer.
Practice runs
AI works alongside your team on real jobs. Clients never see it. It just gets graded.
The bar
You set it — e.g. 98% right, and it earns the job. Below the bar, it keeps practicing.
04 · The report
We didn't model this section — we ran it. The documents are made up (no client data, ever). The grades, the costs, and the mistakes are real, scored case by case on the Recursiv platform.
Fig. 3a — The tape · every graded case from the live run, in order
Reading the misses: the sorting miss picked the right folder but garbled its reply format. Two typing misses were the grader cutting the answer off mid-number; one was real — a parenthesized loss, (12,340), read wrong. Every one of these is exactly what the grading exists to catch.
Fig. 3b — The scorecard · bar: 98% to run alone, 90% to run with review
| The work | Right | Handed back | Cost each | Verdict |
|---|---|---|---|---|
| Sorting client documents20 docs: W-2 · 1099 · K-1 · receipts · traps | 95.0% | 0% | $0.005/doc | On, with review |
| Typing the numbers in20 boxes, graded one box at a time | 85.0% | 10% | $0.005/box | Keep practicing |
| "Where's my return?" replies & doc chasingright = sent without edits | — | — | — | Measured in pilot |
| First draft of a 1040standard complexity | — | — | — | Measured in pilot |
| K-1 tie-outallocations vs. source | — | — | — | Measured in pilot |
| Tax research memo, first draftright = no real error on review | — | — | — | Measured in pilot |
Right = matches the answer key · Handed back = gave it to a person instead of guessing. Neither task cleared the 98% bar on day one — that's the point of grading before trusting. The pilot fills in the remaining rows on your work.
The point: nothing on this page asks for trust. Every row either has a grade or says it doesn't have one yet. A pilot runs thousands of cases on your firm's own work and re-grades every month.
05 · Your data
Recursiv runs on your own servers. The AI comes to the data — the data doesn't travel. Every action it takes is logged for your WISP.
Fig. 4 — Where everything lives. Default: everything inside the green line. Outside AI services are a choice you make, behind a consent gate — never a default.
06 · The path
Who does what, hour by hour. Output: your version of Fig. 2 — real hours.
≈ 2 weeks
AI works alongside your team on real jobs. Clients never see it.
≈ 4–6 weeks
Your version of Fig. 3. What passed, what didn't — against your bar.
End of practice cycle
Behind review gates. Everything else keeps practicing until the numbers earn it.
Ongoing
What's real here: Fig. 3 rows 1–2 are an actual graded run (40 cases, July 29 2026) on made-up tax documents — no client data. Fig. 2 is a representative model of a 16-person firm, and the green shares there are estimates until a pilot measures them. Nothing here is legal or tax advice. The engagement replaces every estimate with a measured number from your firm.
The next step
A 6–8 week pilot: we map your hours, AI practices on your actual work — clients never see it — and you get this exact report with measured numbers. If the grades don't earn it, you switch nothing on.