The Golden Dataset: What Your Firm Owes an AI Vendor Before Day One

Anirudh Badam
Co-Founder and CAIO, Adopt AI10 August 2026

You have decided to bring an AI vendor in to automate part of your workflow. Statement processing, tax prep, reconciliation, whatever the use case is. The contract is signed.
Now comes the part most firms underestimate. Before the vendor can build anything worth using, your firm has to hand over a golden dataset.
Not a handful of clean samples pulled together in an afternoon. The set of real examples that teaches the build what your firm's correct output actually looks like, assembled the way your team already assembles it.
The vendor was not in your meetings. Your team was.

The short answer
A golden dataset is your own completed work, paired with the raw inputs it came from, validated by someone with signing authority. It is what the agents get built and tuned against, and it is the answer key their output gets graded against.
Three things determine whether yours is any good:
- Variance coverage, weighted by client volume, not by segment count. 75% coverage is a legitimate place to start a pilot. 90% is the number to reach before this touches production work at scale. 100% is not available to anyone.
- Both sides of every example. The messy input exactly as it arrived from the client, and the finished output that is correct.
- A documented reason each output is correct, not just a note that it is.
Get this right and the engagement starts on solid ground. Get it wrong and no amount of vendor talent fixes a build shaped around the wrong picture of your firm.

Adoption starts underwater
AI capability is moving on a curve that compounds. Firm buying moves on a straight line, because tools get purchased at a steady, predictable pace. Real adoption lags even that. Plenty of firms are sitting on paid licenses and stalled pilots, which means their effective adoption is running below zero before it ever starts climbing.
The numbers on this are not encouraging. Gartner put the share of enterprises that had scaled AI beyond the pilot stage at 9%. MIT's review of 300 enterprise generative AI deployments found 95% delivered no measurable P&L impact.
The common thread in both is not model quality. It is a mismatch between what the vendor built against and how the business actually runs.
That mismatch has a specific, addressable cause, and it sits on your side of the table.
Every vendor has an accuracy slide
"99% precision." "Trained on thousands of returns." None of it means anything until you ask two questions:
- What exactly was it built and tested against?
- Is any of that relevant to your firm's own structure, positions, and house standards?
A number produced against someone else's client mix, someone else's document conventions, and someone else's definition of correct tells you the technology functions. It tells you nothing about your engagements. Your firm's data is the only instrument that answers the question you are actually asking.
One clarification worth making, because the phrase "training data" causes trouble here. In a well-designed accounting build, the model is not memorizing your numbers and then reciting them back. The dataset is used to design the workflow, tune the extraction and calculation logic, and grade the finished output against a known-correct answer. If a vendor describes your client files as material a model will learn figures from and later reproduce, that is a different architecture with a different risk profile, and it is worth a direct conversation before a single file moves.
Start by writing down your client archetypes
Before you pull a single client record, write down the archetypes the build needs to handle. Not "small business" and "individual." Actual categories.
| Dimension | Why it changes what "correct" means |
|---|---|
| Entity type | Sole proprietor, S-corp, C-corp, partnership, and which flavor of LLC. Each carries different schedules, different elections, different failure modes |
| Filing footprint | Single-state or multi-state, domestic-only or foreign income involved. Apportionment and international reporting are entire bodies of logic |
| Industry | A construction client's books look nothing like a SaaS client's |
| Accounting method | Cash or accrual, plus anyone who switched mid-year |
| Size tier | A $2M business generates different volume and different exceptions than a $200M one |
The build only learns what you show it. Weight the set heavily toward single-member LLC Schedule Cs and it will be confident and wrong the first time it meets a partnership's guaranteed payments.
Document which archetype each example represents as you assemble the set. It is the cheapest insurance in this entire process, and it is the thing nobody does until the second attempt.
This matters most where your practice is heaviest. If you lean into multi-state manufacturing clients, or partnerships with complex allocations, make sure those are well represented from the start rather than added after the build already underperforms on them.
The number that matters: variance coverage
Here is the framework worth using while you build. Map your firm's realistic client variance, then measure what share of it the set actually covers.
75% is a legitimate place to start. A set at 75% coverage teaches the common cases and gets you through an initial pilot. That is enough to prove the approach works on your work.
But 75% also means one in four of your real clients falls outside anything the build has seen. In tax work that is not a rounding error. That is the client with a basis limitation on their K-1, or the multi-state apportionment scenario nobody included, arriving untrained in a live engagement. You find out in front of a client instead of during the pilot.
90% is the number to build toward before this touches production work at scale. Getting there is not about adding more of the same data. It is about deliberately hunting the variance you are missing: amended returns, unusual elections, clients who changed accounting method mid-year, multi-entity ownership structures. Individually rare. Across a full season, they add up to real exposure.
100% is not realistic for anyone. Tax law changes every year and new structures show up. What matters is knowing honestly where the remaining gap sits, and deciding in advance what happens when a live case lands in it. Flag for human review, or escalate. Never guess.

How to actually calculate the percentage
Saying "aim for 90%" is easy. Measuring it takes work, though not complicated work.
This is not a new idea either. It is stratified sampling, the standard statistical method for making sure a dataset represents every subgroup in a population rather than just the average case. Data scientists have used it for decades precisely to stop evaluation sets from skewing toward whatever was easiest to collect.
The step people skip is weighting by client volume instead of by segment count.
Say you have identified 40 segments and your set covers 30 of them. That looks like 75%. But if those 30 segments hold only 600 of your 1,200 active clients, your real coverage is 50%. Same set, two answers, and only the client-weighted one is true.
How many examples does a segment need before it counts as covered? Practice in few-shot classification and fine-tuning generally puts the floor at two to four examples per category, enough for the pattern to appear more than once. Below that you are not covering the segment, you are hoping it generalizes from an anecdote. Go well above the floor for segments carrying real complexity.
Volume tells you what is common. It does not tell you what is dangerous to get wrong. A segment might be 3% of client volume and carry outsized risk, a multi-entity structure with related-party transactions being the obvious example. Flag those separately and cover them regardless of volume share.
That leaves you two numbers to track, and you need both:
- Coverage by client volume
- Coverage of your flagged high-risk segments
A firm at 85% volume coverage with every high-risk segment included is in better shape than a firm at 92% that skipped its three messiest client types because they were rare.
Both sides of the file belong in the set
A golden dataset is not just the answer key. It is the raw material plus the answer key. For an accounting or tax build, that means both halves of every example.
Input files. Bank statements, prior-year processed records, source documents, preparer notes, anything that would land on a staff accountant's desk at the start of the engagement.
Output files. The processed statement, the return, the workpaper, whatever the correct finished product looks like.
Hand over only the clean finished output and the build never learns what a real starting point looks like. Hand over only the input and there is nothing to establish the correct answer from.
One rule here matters more than it sounds like it should. Hand over input files exactly as they arrived from the client. Do not clean them up first.
The instinct is understandable. A statement with inconsistent formatting, missing memo fields, or a handwritten note scrawled in the margin feels like something to tidy before anyone outside the firm sees it. Resist it. If your team cleans the inputs before they go into the set, the build is shaped around a version of reality that does not exist. Then it goes live, meets an actual unedited statement, and stumbles on exactly the mess you filtered out.
Raw in, raw out. That is the only version that teaches anything.
What to hand over besides files
The set is necessary and not sufficient. Six things determine whether the vendor can use it.
Name the owners internally. A manager or senior associate who knows the workflow end to end. A CPA with signing authority who can validate that an output is correct. Someone who can hold technical scope with the vendor. Skip any one of these and either the dataset or the build comes up short.
Get NDAs in place before a single file moves. Standard vendor NDAs are often not built for client data confidentiality. Loop in whoever owns data privacy at your firm first, and check whether your engagement letters need updating to permit this kind of sharing. Worth asking early whether the vendor can deploy inside your own environment, because if client data never leaves your infrastructure, the consent question largely stops existing rather than needing an answer.
Use a secure transfer method, not email. A file transfer portal, a shared drive with access controls, or the vendor's own intake process. Ask what they use before you pick.
Schedule a walkthrough call. Do not hand over a folder and assume it explains itself. Walk their team through how your firm organizes documents and what a margin note actually means. One call up front saves weeks of back and forth.
Write down how you would train a new hire, and give them that too. How you would walk an associate through processing a statement in week one, which shortcuts are acceptable and which are not. The dataset shows what correct looks like. This document shows why.
Pick pilot clients deliberately. A mix of entity types, at least one client with genuine complexity, and ideally one case that gave your team trouble last season. A pilot built on five easy clients proves nothing that was ever in doubt.
The checklist
- Assemble the internal team: a practitioner who can sign off on correctness, a manager who knows the workflow end to end, an owner for technical scope with the vendor
- Confirm NDAs and client engagement letters cover sharing client data with this vendor
- Establish a secure transfer method, and do not default to email
- Map your client base by entity type, filing footprint, industry, accounting method, and size tier
- Select pilot clients that represent real variance, not the easiest files on hand
- Gather input files exactly as received, unedited: statements, prior-year records, preparer notes
- Gather the matching output files: the correct finished return, workpaper, or processed statement for each input
- Have a practitioner validate and document why each output is correct, not just that it is
- Confirm at least two to four examples per segment, more where complexity is real
- Calculate coverage as a share of your realistic client variance, weighted by client volume
- Flag your highest-risk segments separately and confirm each is represented regardless of volume share
- Decide in advance what happens when a live case falls outside coverage: flag, escalate, never guess
- Schedule a walkthrough call to explain your file structure before the vendor starts building
- Write an onboarding-style workflow document, as if training a new hire, and share it alongside the dataset
- Set a refresh cadence so the set does not go stale when rules change
Why this outlives the vendor
One argument for doing this properly that rarely gets made: the golden dataset is a firm asset, not a vendor deliverable.
It is the only instrument that lets you compare two vendors on the same work instead of on two different demos. It is the baseline that turns "how accurate is it?" from a sales question into a measurement. It is what makes a second use case cheaper than the first, because the archetype map and the validation logic already exist. And if a vendor relationship ends, the set is what stops you from starting the next evaluation from zero.
Which is a reason to be clear from the outset about who owns the dataset, the documented positions inside it, and the logic built on top of it. The answer should be your firm.
Building this for a real engagement? Book a pilot. We build against your workpapers, your client structures, and your unedited source files, not a demo dataset, and we co-build with your own team so the archetype map and the validation logic stay yours.
Smaller firm, want to test the idea first? Start free, pull two or three finished engagements you already know are correct, and run them through as your own answer key.
Written from delivery experience building agents against firms' own engagements. Coverage targets and per-segment example counts are working guidance rather than measured thresholds, and the right numbers for your firm depend on your client mix and your risk tolerance.





