Will AI Replace Accountants? What Working Inside a Top-30 Firm Actually Showed

Sunil Neurgaonkar
Marketing6 August 2026

Every answer to this question is a forecast. Someone with a model of the future tells you what the profession will look like in ten years, and you are asked to take it on faith because there is nothing else on offer.
This is not that. We ran AI agents against live R&D tax credit engagements inside a top-30 US accounting firm, on real client work rather than a demo dataset, and watched what the agents finished, what they got wrong, and what a human still had to do afterward. What we found was not what either the optimists or the pessimists would have predicted.
Here is the short version, before the argument.
The short answer
No, accounting will not be replaced by AI, and the reason is structural rather than sentimental.
Break a real engagement into hours and it splits roughly into thirds of very different character. About 40% of staff time is calculation and assembly work, and most of that is automatable today. Another 40% is client communication, site visits, and document management, where agents do essentially nothing and there is no credible path to them doing so. The remainder is coordination overhead and review.
What changed on the engagements we ran was not the headcount. It was the composition of the work. The senior person on the file stopped rebuilding a workbook and started reviewing exceptions in it.
So the honest formulation is not "will AI take over accounting jobs." It is: the preparation layer of accounting is automatable now, the judgment and relationship layer is not, and most accounting roles are a blend of the two in a ratio that is about to change.
That is a real answer with real consequences, including uncomfortable ones. Those are further down.
Why the question keeps getting asked wrong
"Will accounting be replaced by AI" treats accounting as a single thing that either survives or does not. Nobody who has actually staffed an engagement thinks that way.
Pull apart a single R&D credit study and you get a list of activities with wildly different exposure:
- Chasing a client for the payroll register: not automatable, and not because of technology
- Pulling that register out of a portal with no API once it arrives: automatable
- Deciding whether a given project meets the four-part test: not automatable
- Applying that decision consistently across 60 employees and three years: automatable
- Explaining to the client why their controller's number is wrong: not automatable, and the highest-value hour in the engagement
- Building the workpaper that proves the number: automatable
- Signing the return: not automatable, ever, for reasons that are legal rather than technical
The word "accountant" covers all seven of those. The forecast question forces you to average them into one answer, which is why the answers are useless. Stop asking whether the profession survives and start asking which of your hours are which.
How we looked at it, and the bar we held it to
The setup. A top-30 US accounting firm. R&D tax credit engagements, chosen deliberately because they are calculation-heavy, high-volume, and staffed by people whose time the firm would rather spend elsewhere. Agents ran inside the firm's own environment, against the firm's own client documents, in the systems the work already lived in, with the firm's own subject matter experts in the loop throughout rather than at a hand-off at the end.
The scope. Live engagements, not a demo dataset. The agents pulled source documents, extracted and categorized the underlying data, built the qualification workpapers, ran the calculations, and drafted the technical memorandum. Then a human reviewed all of it, and we paid attention to what that review cost.
The bar. Not "time saved," which is easy to inflate and impossible to check. One of the firm's own leads put the real standard plainly: what matters is net hours saved after accounting for model review and error corrections, because automation that creates rework negates time savings.
That is a harder bar than most AI claims are built to clear, and it is the one worth holding any vendor to, including us. Everything below is what survived it.
The 40% that automated, and the 40% that did not
The single most useful output of the engagement was not a percentage. It was a clean split of the engagement into three buckets, which turns out to hold well beyond R&D credits.
| Activity | Automatable today | Why |
|---|---|---|
| Retrieving documents from client portals, banks, and ERPs | Yes | Agents operate the interface the way a person does, so systems with no API are in scope rather than blocked |
| Extracting and categorizing source data | Yes | Deterministic once the extraction logic is written and reviewed |
| Building workpapers and tie-outs | Yes | Structural work with a defined right answer |
| Applying a decided position consistently across a population | Yes | This is exactly what machines are for and exactly what humans do worst at hour nine |
| Rolling prior-year positions forward and proving they tie | Yes | Mechanical, and the most common source of multi-year error |
| Drafting the supporting narrative | Mostly | Close enough to need a human editor rather than a human author |
| Deciding the position | No | Professional judgment, and the thing the client is actually paying for |
Notice what is not on the automatable list: anything involving a decision that could be wrong in a defensible-but-arguable way. And notice what is: nearly everything a first-year associate spends their first two years doing.
That asymmetry is the whole story, and it is why "will AI replace accountants" gets a different answer depending on which accountant you ask.
The part that does not fit on a slide
Four things went differently than we expected, and all four cut against the marketing version of this technology.
The tuning cost was substantial and front-loaded. This was not a tool that was switched on. Getting an agent to finish an engagement to a standard a manager would sign required heavy iteration against real client files, with the firm's own subject matter experts putting serious hours into agent design rather than sitting on the sidelines. Anyone selling you a two-week transformation is selling you a demo. The honest framing is that you are building something with a vendor, not buying something from one.
Stable clients and new clients are different problems. Completion rates hold up on multi-year clients where prior-year positions exist to anchor against. First-year clients, clients with a restructure, clients whose documentation habits changed: all materially lower. Any firm evaluating this should ask for completion rates split by client stability, because a blended average hides the number that determines your staffing plan.
Review did not disappear. It changed shape and it got harder to do badly. Review hours came down sharply, but they did not go to zero, and they should not. Reviewing agent-produced work is a different skill from reviewing a junior's work. You are not checking whether someone understood the assignment. You are checking a completed artifact for the specific failure modes of the thing that produced it. Firms that assume their existing review process transfers unchanged are the ones that will get burned.
The ceiling was set by process, not by intelligence. The share of time that is neither cleanly automatable nor obviously human was mostly coordination overhead: waiting on documents, chasing approvals, resolving who owns what. No model fixes that. A firm with a disciplined PBC process gets more out of this than a smarter model would give a firm without one.
The uncomfortable version of the answer
The reassuring framing, and it is the one every vendor including us reaches for, is that AI frees accountants from grunt work so they can do higher-value advisory work. That is true. It is also incomplete, and you should hear the incomplete part from someone selling the technology rather than discovering it later.
Three things are genuinely at risk.
The pure preparation role. A job that consists entirely of assembling data into a workbook is a job that consists entirely of the automatable 40%. That role does not survive in its current form. In practice we have watched this land first not on staff accountants but on offshore preparation teams, which are the actual incumbent being displaced. In one engagement, an offshore team was pulling documents from 15 separate bank portals with scripts and hand-keying them into Excel. That workflow is the most exposed thing in the accounting supply chain, and it is exposed today rather than in five years.
The traditional training ladder. The profession teaches judgment by making juniors do assembly work for two years until pattern recognition sets in. Remove the assembly work and you have removed the curriculum. Nobody has solved this yet, and firms that automate preparation without deliberately rebuilding how juniors learn will find out in about four years that they have a bench of reviewers who never learned what they are reviewing. This is a real risk and it deserves more attention than it gets.
Firms that price by the hour on automatable work. If your revenue on an engagement is a function of hours spent assembling a workbook, and that assembly compresses by 85%, your revenue model has a problem that is independent of whether you adopt the technology. Your competitor adopting it is sufficient.
What is not at risk: the license, the signature, the judgment, the client relationship, and the entire population of accountants whose work already skews toward those. Which, in most firms, is everyone above the second year.
The real question is capacity, not replacement
Here is what makes the replacement framing feel so strange from inside a firm: nobody in this profession has too many people.
The pool of qualified accountants is shrinking. Firms cannot hire the staff they need at the price they can pay. Partners are turning down work, or accepting it and absorbing it personally. Growth is capped by headcount rather than by demand. Against that backdrop, "AI is coming for accounting jobs" describes a labor surplus that does not exist.
The constraint is capacity. Every firm we talk to is choosing between three ways to buy it:
- Hire. The pool is shrinking and compensation is rising. Scales linearly at best.
- Offshore. Management overhead, quality variance, still-linear cost, and in one case we saw, a requirement to get explicit consent from 4,000 clients before client data could go anywhere.
- Automate the preparation layer. Front-loaded effort, no consent problem if the work stays inside your own environment, and it does not scale linearly.
Notice that "do nothing" is not on the list. It has no cost this year and a compounding one after that.
Framed this way the anxiety inverts. The firm that automates preparation is not the firm shedding accountants. It is the firm that can finally say yes to the work it has been declining, with the people it already has.
What this means depending on where you sit
If you are doing the work. Audit your own week against the three buckets above. If more than half your hours are in the automatable column, that is worth knowing now rather than in three years, and the response is not to leave the profession. It is to move deliberately toward the review, judgment, and client side, which is where the profession has always paid best anyway. The skill that appreciates fastest is being the person who can review machine-produced work rigorously, because almost nobody has it yet.
If you run a team. Your review process is the thing that needs redesigning, not your headcount. Ask what your reviewers check for, and whether those checks catch the failure modes of an agent rather than the failure modes of a tired junior. Then ask how your first and second years will learn judgment once they stop building workbooks, and write down an actual answer.
If you run the firm. The decision is not whether AI replaces accountants. It is whether your capacity ceiling is set by hiring in a market where you cannot hire. Run the math on the 40%, on your own engagements, with your own realization rates, before you run it on a vendor's slide.
How to tell whether a tool actually does the work
The market is full of things that answer questions about accounting and very few that do accounting. Six questions that separate them, drawn from what actually mattered in delivery:
- Does it operate my systems, or does it need an API? Most accounting work lives in portals, legacy systems, and client environments with nothing to integrate against. A tool that requires clean APIs cannot reach the work.
- Where does the number come from? If a language model reads a figure out of a document and types it into a workpaper, you have a plausible number with no provenance. The safer design is that the model writes the conversion logic, a person reviews that logic, and the logic moves the value. The model never touches the number.
- Can I trace every populated cell to its source? If a reviewer cannot click a cell and see where it came from, they will re-derive it, and you have added work rather than removed it.
- What checks it before it reaches me? Ask specifically. On our engagements a checker agent runs a 100-point validation with audit sampling on key values in a maker-checker loop, before a human ever sees the output.
- What is the completion rate on clients like mine, split by stability? Not the blended average. Not the best case.
- Where does my client data live? If the answer requires client consent to send data somewhere, you have just added a consent project to your engagement. Deploying inside your own environment removes the question rather than answering it.
If a vendor cannot answer all six specifically, you are looking at a demo rather than a capability.
<div class="cta-block">
Want the version of this measured on your own engagements? Book a pilot. We build against your workpapers and your client structures, not a demo dataset, and the pilot is scoped so the answer is real before the contract is.
Smaller firm, want to see it yourself? Start free and run one live engagement through it.
</div>Related reading
- Form 5471: Who Must File, Every Schedule, and the Errors That Trigger Penalties
- Transfer Pricing Documentation: A Practitioner's Guide to Scoping and Defending It
This post describes what we observed on R&D tax credit engagements at one firm. Results on other workflows, and at firms with different processes and client mixes, will differ.




