The Modern CPA Firm Tech Stack, and Where Agents Fit

Deepak Anchala
Co-Founder and CEO, Adopt AI28 July 2026
Ask a partner what their firm's tech stack is and you will usually get a list of logins. Tax software, the ledger, the portal, the workflow tool somebody championed in 2022, a document system that half the staff bypass, and Excel, which is where the work actually happens.
That list is not a stack. It is an accumulation. Most of it was bought one product at a time, each to solve one visible problem, and the seams between the products were never anybody's purchase.
This guide is about the seams. It maps each layer of a modern firm's stack, says plainly what that layer does and what it leaves for a person to do by hand, shows how the shape changes as a firm grows from one person to several hundred, and then works through the decision most firms are facing right now: whether to buy AI capability, build it, or build it with someone.
A caveat up front, because it sets the tone for everything below. If you are reading this while sitting on fifteen demo requests, the useful thing this page can do is not introduce a sixteenth product category. It is to give you a way to tell which of your problems are actually tooling problems, and which are process problems that no purchase will fix.
Who this is for: firm-side practitioners and owners, at any size from a solo practice to a multi-office firm. The layer map holds regardless of size. The buying advice changes considerably by size, so the sizing section is worth reading for your own band rather than in general.
What a CPA firm tech stack actually is
A firm tech stack is the set of systems that carry client work from intake to delivery, plus the systems that keep the firm itself running. It is useful to think of it in nine layers rather than as a list of vendors, because layers are what you can reason about. Vendors move between layers, merge, and rebrand. The layers do not change.
The right-hand column is the one that matters, and it is the one vendor comparison pages leave out.
| Layer | What it does | Common systems | What it leaves manual |
|---|---|---|---|
| 1. Practice management and workflow | Job tracking, due dates, capacity planning, staff assignment, time and billing | Karbon, Canopy, Aiwyn, Jetpack Workflow, Firm360 | The work itself. It knows a return is due and who owns it, not how to prepare it |
| 2. Tax compliance | Return preparation, calculation, e-filing, organizers | CCH Axcess, UltraTax, GoSystem, Lacerte, ProConnect, Drake | Getting source data into the input screens, and building the workpapers behind the numbers |
| 3. Audit and assurance | Engagement structure, trial balance, workpaper binders, sign-off trails, PBC tracking | CaseWare, Cloud Audit Suite, TeamMate, Suralink, Fieldguide | Tie-outs, populating the workpaper, chasing the client for the list |
| 4. Client accounting and ledger | The books themselves: GL, AP, AR, multi-entity | QuickBooks Online, Xero, Sage Intacct, NetSuite, Dynamics 365 | Cleanup, categorization exceptions, anything consolidated outside the system |
| 5. Close and reconciliation management | Close checklists, task sequencing, reconciliation status, variance tracking |
Two observations follow from the table, and both are uncomfortable.
First, every layer here has good products in it. The dissatisfaction most firms feel is not because layers 1 through 6 are badly served. They are well served. Practice management genuinely tracks work. Tax software genuinely computes returns. Close platforms genuinely enforce a sequence.
Second, the hours do not sit inside any of those layers. They sit in layer 7, and in the movement between layers: pulling twelve months of statements out of a bank portal, getting a client's messy trial balance into a shape the audit binder will accept, keying extracted figures into tax input screens, rebuilding last year's workpaper because the source data arrived in a different format again.
Nobody sells that. It has no product category, so it never appears on a stack diagram, so it is never on the agenda when a firm reviews its technology spend. It is just what staff do.
The seam problem, stated precisely
Here is the pattern that explains why firms with excellent tooling still have a preparation bottleneck.
Every system in the table above is designed around a boundary. It accepts data in a defined shape, does something valuable with it, and emits data in a defined shape. Within the boundary, software. On the boundary, a person.
For firm-side work, those boundaries are unusually numerous and unusually hostile, for three reasons that are structural rather than fixable by buying better products.
Your clients own half the stack, and you do not choose it. A firm with 300 clients has some number of distinct ledger systems, bank portals, payroll providers, and filing habits among them. You inherit all of it. Every promise of a clean integration assumes control over both endpoints, and you have control over one.
A large share of the source systems have no usable API. Bank portals, county and state filing sites, insurance carrier systems, client-side legacy ERPs, and a great many industry-specific platforms simply do not offer programmatic access to the data you need. This is not a temporary condition. For anything built on API integration, those systems are permanently out of scope, which is why an integration-led stack strategy stalls at exactly the point where the manual hours are.
The output has to be defensible, not just correct. A number that appears in a workpaper without a traceable path back to its source costs more in review than it saved in preparation. This is the constraint that most general-purpose tooling fails, and it is why "it produced the right answer" is not sufficient evidence for this profession.
We treat the layer model of the work itself, and the nine specific reasons automation projects fail inside real firms, at length in the accounting automation guide. This page stays on the stack question: given those constraints, what should you actually own?
How the stack changes with firm size
The layer map is constant. What varies is which layers you buy, who owns them, and what breaks first. Find your band.
Solo through five people
Typical stack: tax software, one or two client ledgers, a shared drive or portal, Outlook, Excel, a payments processor. Practice management is a spreadsheet, or the owner's memory.
What works: almost no coordination overhead, and no committee between noticing a problem and fixing it. A solo practitioner can adopt something useful on a Tuesday afternoon. That speed is a genuine structural advantage over large firms, and it is routinely underrated.
What breaks first: the owner is the single point of failure for every layer at once, including the ones that are not accounting. Capacity is capped by one person's hours, so the constraint is never software features, it is whether a tool removes hours from the owner's week in the first fortnight. Anything with a learning curve longer than the pain it relieves will be abandoned.
The buying rule at this size: only add something you can evaluate on your own real work, this week, without a call. If a tool cannot be tried against a job already sitting on your desk, it is priced in your time before it is priced in dollars.
Five to twenty people
Typical stack: everything above, plus practice management, a document portal with client-facing requests, e-signature, and usually a first attempt at standardized workpaper templates.
What works: this is the band where standardization first pays for itself. Templates, naming conventions, and a real workflow tool convert the owner's judgment into something staff can execute.
What breaks first: nobody owns the stack. Tools get adopted by the person who felt the pain, adjacent staff opt out quietly, and within a year a client's documents live in four places with no authoritative copy. The visible symptom is usually a partner asking why nobody can find last year's file.
The buying rule at this size: before any new tool, name the person accountable for it after go-live, and name what it replaces. If the answer to the second question is "nothing, it adds visibility," you are buying overhead, which is occasionally the right call and should at least be a deliberate one.
Twenty to two hundred people
Typical stack: all of the above, plus an audit or assurance suite, a real document management system with retention rules, security tooling that survives a client questionnaire, offshore or outsourced preparation capacity in many cases, and multiple overlapping tools inherited from whoever championed them.
What works: specialization. There are enough people to have a tax stack, an audit stack, and a client accounting stack, each fit for purpose, and someone whose job includes firm operations.
What breaks first: overlap and seams, in that order. Three systems can each hold a version of a client's trial balance. The offshore team develops its own process outside every system you paid for, typically a set of scripts and a shared workbook, and that process becomes load-bearing and undocumented. This is also the band where the preparation bottleneck becomes a growth ceiling rather than an annoyance, because client load is now limited by how many engagements your seniors can carry through preparation, not by demand.
The buying rule at this size: audit before you add. The stack inventory exercise below usually finds two consolidation opportunities that free more capacity than the next purchase would.
Two hundred and up, multi-office
Typical stack: the above, times the number of offices and the number of firms acquired, plus IT and InfoSec functions with veto authority, formal vendor review, and increasingly a data or AI team of some size.
What works: the firm can actually run a controlled evaluation, absorb a security review, and deploy something consistently once decided.
What breaks first: consolidation politics and acquisition sprawl. Each acquired firm arrives with its own stack, its own conventions, and partners with strong opinions about both. Technology decisions become organizational decisions, and the honest constraint on adoption is rarely the technology.
The buying rule at this size: run it as a scoped pilot on one workflow in one office with written exit criteria, and settle the ownership question before the security question. If the firm is owned by a private equity platform, the build-versus-buy decision may not sit with the firm at all, which is worth establishing early rather than discovering three months in.
Where agents fit, and where they do not
The category name is doing you no favors here, so let us be concrete about the claim.
Agents in this context are software that operates your existing systems the way a person does, through the interface when there is no API, in order to produce a work product: retrieve the statements, extract the figures, build the reconciliation or the workpaper or the statement set, tie it out, and hand a person the exceptions.
That is a claim about layer 7 and the seams, which is the part of the stack nothing else addresses. It is not a claim about replacing layers 1 through 6. Your tax software still computes the return. Your audit suite still holds the binder and the sign-offs. Your practice management still tracks the job. What changes is who produces the input those systems have always required a person to produce.
Three properties determine whether this is real for your firm, and they are the ones to interrogate in any evaluation.
Does it reach systems with no API? If the answer involves an integration roadmap, then the systems where your hours are concentrated are out of scope, and the pilot will succeed on the easy 30% of your work. Ask with your actual system list, including the client-side ones and the bank portals, and ask to see it done on one of them.
Is the output traceable? A populated workpaper should carry formulas that point back to the source tab, and agent-populated cells should be distinguishable from human ones. A reviewer who has to re-derive the numbers is doing preparation with extra steps. This is also the honest answer to the hallucination objection: the strongest available design has the model write the conversion logic that a person can read and check, rather than having the model report figures directly, so the numbers are produced by reviewable code rather than by a model's assertion.
Does it produce exceptions, or output? The valuable artifact is a short queue of items needing judgment, next to a completed artifact. A tool that returns 100% of the work for checking has moved the labor rather than removed it.
Where agents do not help, stated plainly, because a guide that claims everything is worth nothing:
- Deciding positions. Method selection, estimates, materiality, treatment. These are the professional judgment the engagement exists to provide.
- Client communication and chasing. Structurally a large share of staff time on many engagements, and structurally not automatable.
- Sign-off. Someone with a license takes responsibility. That does not move.
- Genuinely one-off work. Volume and repetition are what pay back the setup effort. A workflow you run twice a year for one client is a bad first candidate no matter how painful it is.
- A broken process. If three staff each prepare the same schedule differently and none of it is documented, the first output of any automation effort will be a documented process, and you should budget for that as work rather than discover it as a delay.
Build, buy, or co-build
This is the live decision, and it deserves more than a vendor's answer. Firms are choosing between three paths right now, and all three are defensible under different conditions.
| Buy a product | Build in-house | Co-build | |
|---|---|---|---|
| Who defines the workflow logic | The vendor, generalized across customers | Your people | Your people |
| Who writes the software | The vendor | Your engineers, if you have them | Vendor engineers working alongside your staff |
| Time to first working output | Fast, if your workflow matches the product | Slow, and usually slower than estimated | Fast, but requires real time from your SMEs |
| What you own at the end | A subscription | Everything, including the maintenance | The agents, skills, and decision logic as your IP |
| Fit to a specialty workflow | Poor if the workflow is unusual | Excellent | Excellent |
| Key risk | The 20% of your work that does not fit the product is still manual |
When building in-house is genuinely the right call
Some firms should build, and pretending otherwise is not credible. Build when all four of these are true:
- The workflow is specific to how your firm competes, not just specific to your firm. A proprietary approach to a specialty practice qualifies. A slightly unusual reconciliation format does not.
- You can staff it permanently, with more than one person. Two engineers who can maintain it, not one enthusiastic senior manager doing it on evenings.
- The surface it depends on is stable. Building against your own ledger is one thing. Building against 15 client bank portals that each redesign without notice is an ongoing maintenance obligation, not a project.
- You have somewhere to put the controls. Version history, an audit trail, an exception queue, access control, and a review workflow. Working output without these is a demonstration, not something you can put client work through.
How in-house builds actually fail
Not on capability. Firms build impressive things. They fail on four patterns worth naming, because each one is avoidable if you plan for it:
- The single champion. One talented person builds something that works, and when they change roles the firm discovers nobody else can modify it. This is the same failure mode as the 2019 Excel macro, at higher stakes.
- Demo to production is most of the work. Getting a workflow to run correctly on a well-behaved client is perhaps a fifth of the effort. The rest is the long tail of client variation, error handling, provenance, and the review process, and the last of those is process design rather than engineering.
- The build outruns its own assumptions. The underlying model layer changes faster than a firm's build cycle. Something architected around one generation's constraints frequently becomes the wrong shape within a year.
- No process change alongside it. If reviewers still check every figure by hand, the tool has produced no capacity regardless of how well it runs. Measured savings live in the review process as much as in the preparation.
What co-build actually means
Co-build is the middle path, and for a firm that has already started building internally it is usually the honest recommendation rather than a compromise.
Your subject matter experts define the logic, because they are the only people who know how the work is actually done at your firm, including the exceptions that are not written down anywhere. Vendor engineers build alongside them. Your staff are deliberately brought up the curve during the build, so the capability stays after the engagement. The agents, skills, and decision logic are your intellectual property, which means you are not one vendor acquisition away from starting over.
The practical significance for a firm with an internal AI effort underway: co-build makes that effort an asset in the conversation rather than a reason to end it. You are not choosing between your team and a vendor. You are choosing how fast your team gets to something in production, and how much of the undifferentiated work they have to do themselves.
The cost of this path is honest and worth stating: it requires real time from your best people, typically tens of hours each from the SMEs who understand the workflow. Firms that will not commit that time should buy a standard product instead and accept the fit compromise, because a co-build without domain input produces a generic tool at a specialized price.
Auditing the stack you already have
Before adding anything, run this. It takes a couple of hours and it consistently changes what firms buy next. The point is not the inventory, it is the last two columns.
| System | Layer | Who owns it internally | What it replaced | Hours it removes per month | Hours of manual work it creates |
|---|
Fill a row per system, then look for four specific patterns:
Systems with no internal owner. These are the ones that will be renewed automatically and used by nobody. Every firm above ten people has at least one.
Systems whose value is visibility rather than output. Legitimate, and worth knowing about. A firm that believes it bought production capability and actually bought tracking will keep buying tracking and keep being disappointed.
Duplicate homes for the same artifact. Two systems holding a client's documents is a filing tax on every engagement, paid in staff time and in the risk of working from the wrong version.
The right-hand column, which is the finding. Almost every system in a firm creates some manual work: exporting, reformatting, keying between systems, filing. Add that column up. That total is your seam cost, it appears in no budget, and it is usually larger than the next license you were considering.
Then ask the ownership question about everything on the list: if we stopped paying this vendor tomorrow, what do we keep? For data, the answer should be "all of it, in a usable format." For process logic, most firms discover the answer is "nothing," which is worth knowing before it matters.
Ten questions to ask before anything joins the stack
- Which layer of the work does this perform, and which layer does it only organize? Get a straight answer, then verify it against the demo you were shown.
- Show me this running on my system list, including the ones with no API and the client-side ones. Not a roadmap.
- What does the output look like to a reviewer? Can they see where each number came from without re-deriving it?
- What does it produce when it is not confident? An exception queue is the right answer. A silent best guess is a disqualifier.
- Where does client data live and move? Our environment, your cloud, or a third-party model provider? This determines whether we need client consent, which determines whether this is possible at all.
- What tuning effort does this take on our workflow and client mix, measured in our hours and your hours, before it performs at the level in the demo?
- Who owns the workflow logic we build inside this? If we leave, what comes with us?
- What breaks it? A UI change on a client portal, a new document format, a client with an unusual structure. What happens then, and who fixes it?
- What is our baseline, and how will we measure against it? If neither party can state today's hours, no savings claim afterwards will be checkable.
- What process changes on our side for this to produce capacity rather than just output? If the answer is "none," the savings will not appear.
Question 6 and question 9 are the ones vendors are least prepared for, and the answers separate a real evaluation from a good demo.
What to measure once something is in the stack
Firm technology decisions get made on feature lists and reviewed on invoices, with nothing measured in between. Four numbers, tracked from before you start:
Net hours saved, after review and rework. Gross preparation time removed, minus review time added, minus time spent fixing output. This is the only figure that corresponds to capacity, and it is the standard a buyer in this profession should apply, because automation that creates rework negates its own time savings.
Completion rate, split by client stability. The share of an artifact produced without human intervention, reported separately for stable multi-year clients and for first-year or unusual ones. A single blended number hides the thing you most need to know, which is how much of your book this actually covers.
Exception rate and exception mix. How often a person is pulled in, and for what. A falling exception rate is progress. A stable rate with a changing mix means you are finding new client variation, which is normal early and a problem late.
Cycle time from documents in hand to reviewable artifact. The number a partner feels, and the one that determines whether you can take on the next engagement in January.
Treat gross time savings, accuracy figures with no tuning cost attached, and any percentage with no stated baseline as marketing rather than evidence, including ours.
A short version, if you are deciding this quarter
- Your stack probably has good products in every layer and an unowned gap between them. The gap is where the hours are.
- Inventory before you buy. The manual work your current systems create is usually a larger number than the next purchase would address.
- Standard high-volume workflow, standard product. Specialized workflow that defines how you compete, build or co-build.
- Build in-house only if you can staff it permanently with more than one person, and only where you control the surface it depends on.
- If an internal AI effort is already underway, co-build is the path that keeps it and shortens it, and it leaves the logic as your IP.
- Start with the workflow that has the highest volume and the lowest variance, on your most stable clients. Baseline the hours first.
- Insist on traceable output and an exception queue. Everything else is negotiable.
Product names above are cited as category exemplars to make the layer map concrete. Inclusion is not an endorsement and the lists are not exhaustive. Category boundaries move as vendors add adjacent capability, so verify current scope against your own workflow rather than against this table.
