Skip to main content
FIELD REPORT · AI TAX SEASON

The AI Stack That Got One Firm Through Tax Season Without Hiring

A real 2026 tax-season case: Blue J + Intuit Tax Advisor + Karbon + Dext + Claude — what worked, what did not.

PUBLISHED
May 13, 2026
READ TIME
8 MIN
AUTHOR
ONE FREQUENCY
KEY FACTS
Topic
AI tax season, tax season productivity, Blue J tax
Industry
accountants
Published
May 13, 2026
Read time
8 min
Word count
1,487

Tax season is the moment of truth for every CPA firm's AI stack. The vendor that looked promising in October either compresses a 14-hour return into 6 hours in March, or it does not. The firm that experimented through the summer either holds revenue per partner constant against a 12% drop in available hours, or it watches realization slide 8 points and call it "the price of growth."

This article is the real 2026 tax-season case: what one 4-partner firm actually ran, what worked, what did not, and what the comparable firm should run next season. The names of the tools are real; the numbers are composited from work with three firms in the $1.4M–$3.2M revenue band.

The firm

  • Profile. 4 partners, 1 manager, 5 staff, 2 admin. $2.1M revenue. 640 1040 clients, 115 entity clients, 22 CAS clients.
  • Stack at start of season. Karbon at practice, mixed QuickBooks Online and Xero at ledger, UltraTax CS for tax prep, Dext on document capture, Aiwyn on proposal-to-payment.
  • Season constraint. One staff senior on maternity leave through April 15. Effective head count down 12% against last season.

The owner's question in November: can we run the same season at lower head count without sliding realization?

The AI stack the firm ran

Six layers, named vendors, real configurations.

Layer 1 — Black Ore on document extraction into UltraTax CS

Black Ore took the inbound 1040 source documents (W-2, 1099, K-1, 1098, brokerage statements), classified them, extracted the key fields, and wrote them into the corresponding UltraTax CS input screen. 1040 prep time on returns with brokerage and K-1 complexity dropped from 3.2 hours to 1.4 hours.

What worked: brokerage statement extraction was strong on Schwab, Fidelity, and Vanguard formats. K-1 extraction was strong on standard PE-fund formats.

What did not: K-1 extraction on partnerships with complex pass-through entities and multi-state allocations was inconsistent. The senior reviewed every K-1 line, which compressed the time-savings on those returns to maybe 30% rather than 60%.

Layer 2 — Intuit Tax Advisor on tax planning

Intuit Tax Advisor ran the year-end planning scenarios for the 22 CAS clients and a subset of 90 1040 clients with planning engagements. Scenario-generation time dropped from 35 minutes per client to under 10.

What worked: the standard "Roth conversion vs. ordinary income deferral" scenario, the QBI optimization scenario, and the year-end estimated-tax-payment recommendation.

What did not: state-specific tax-planning scenarios were weak. The firm dropped back to manual analysis for multi-state clients.

Layer 3 — Karbon AI on email triage and PBC chasing

Karbon AI ran the front of the engagement — classifying inbound emails, drafting standard responses for partner review, and chasing missing PBC documents. Email volume to partner inboxes dropped 45%; PBC follow-up cycle compressed from 8–12 days to 3–5.

What worked: the inbound classifier on "where's my refund?" and "when do you need the rest of my documents?" — the high-volume low-complexity emails that historically consumed partner attention.

What did not: complex client questions ("can I deduct this rental loss?") were routed to partner queue rather than answered directly, which was the right outcome. The AI did not pretend it could answer those.

Layer 4 — Dext on receipt and bill capture

Dext continued doing what it had done all year — capturing receipts and bills for the CAS clients. No tax-season-specific configuration. 100% steady-state.

Layer 5 — Materia AI on tax research

Materia AI handled the embedded research. "Does the new Rev. Proc. 2025-32 SALT-cap workaround apply to this client?" "What's the basis treatment on this S-corp shareholder distribution?" Research time per query dropped from 25 minutes to under 8.

What worked: the citation-anchored answers came with the underlying authority, which the senior validated before relying on the answer.

What did not: novel questions involving 2025-issued guidance with limited interpretation history occasionally produced overconfident answers. Senior validation caught these; an unsupervised junior would not have.

Layer 6 — Claude Enterprise on review-memo and client-communication drafting

Claude Enterprise drafted the partner review memo, the client transmittal letter, and the standard client emails. The textbook transcription-drafting workflow.

What worked: a 35-minute partner review-memo task compressed to 8 minutes (5 minutes for Claude to draft, 3 for the partner to edit).

What did not: anything requiring novel tax analysis. Claude drafted; the partner did the analysis.

The numbers

Real metrics from the 2026 season (Feb 1 — Apr 15).

  • Returns processed. 612 1040s, 108 entity returns. Within 4% of prior-year volume despite 12% lower head count.
  • Average 1040 prep time. 2.1 hours, down from 3.0 hours prior year. 30% compression.
  • Average entity prep time. 5.4 hours, down from 7.8 hours prior year. 31% compression.
  • Realization. 92% against 88% prior year. 4 points of recovery.
  • Partner hours per week, peak weeks. 58 hours against 71 hours prior year. 18% compression.
  • Staff overtime cost. $42,000 against $71,000 prior year budget.
  • Total AI-stack cost for the season. $28,400 in incremental vendor spend.

Net P&L impact for the season: roughly $310,000 of recovered margin against $28,400 of incremental cost. Plus the avoided hire of a contract preparer.

What the firm would do differently next season

Three calls from the post-mortem.

  1. Wire Aiwyn deeper on engagement letters and renewals. Aiwyn ran proposal-to-payment but did not run the engagement-letter-renewal sequence. Next season the firm runs Aiwyn end-to-end including renewal communications.

  2. Pilot a second tax-research model alongside Materia AI. Materia AI was strong; a second tool for cross-validation on novel questions would have caught the two overconfident answers earlier.

  3. Start the staff training in October, not January. Staff hit peak productivity on the AI stack in mid-March. Earlier training would have moved that to mid-February and unlocked another 4–5 points of season-wide productivity.

Lessons that generalize

Six things any comparable firm should do.

  1. Pick one tax-prep-extraction tool. Black Ore or a comparable extraction layer is the highest-leverage single tool. Pick one and configure it well rather than spreading across two.

  2. Wire the practice-management layer first. Karbon AI or Canopy at the front compounds every downstream workflow. Get the front clean.

  3. Run the embedded-research tool. Materia AI or a comparable research layer returns 15–25 minutes per query. On a 14-week season that compounds to 80–120 partner-hours.

  4. Use enterprise-tier drafting AI for every memo and client email. Claude Enterprise or ChatGPT Enterprise; never consumer tier.

  5. Document the billable-time-leakage before and after. Without baseline measurement, the partner cannot defend the AI spend at month nine when budget conversations start.

  6. Build the partner-review gates. Every AI output touching a return goes through a partner review gate before signing. SSTS is non-negotiable.

The full season playbook lives in the 2026 firm playbook and the compliance posture in the AICPA, IRS Pub 4557, and FTC Safeguards article.

What good looks like

Four metrics every CPA owner should hold next season's AI stack to.

  • Returns per FTE per season. Floor 110–135. Target 145–165.
  • Average 1040 prep time. Floor 2.8–3.4 hours. Target under 2.2 hours.
  • Realization through busy season. Floor 84–88%. Target 92%+.
  • Partner peak-week hours. Floor 65–75. Target under 60.

These feed the CPA AI ROI walkthrough.

FAQ

Q: Will this work on a firm running CCH Axcess instead of UltraTax CS? A: Yes. Black Ore integrates with CCH Axcess as well as UltraTax CS and Lacerte. Materia AI is system-agnostic. Karbon AI sits above the tax-prep layer.

Q: What about Drake or ProConnect? A: Drake integration with Black Ore was added in late 2025 and is still maturing. ProConnect (Intuit) gets the Intuit Assist integration natively. Configure carefully and validate in shadow mode before live traffic.

Q: How do I handle the intake-automation layer for new clients during season? A: Treat new-client onboarding the same way mid-season as off-season — Karbon AI or Canopy at intake, Aiwyn on engagement letters, Black Ore on document extraction. See the onboarding automation walkthrough.

Q: What about lead-response-time on inbound during season? A: Same playbook as off-season. The marketing engine should run continuously. See the marketing and lead generation article.

Q: Can solo practitioners run a comparable stack? A: Yes, at lower price points. Pixie or Karbon Solo plus Black Ore solo tier plus Materia AI plus Claude or ChatGPT Enterprise lands at $6k–$11k/year. The ROI math holds.

Q: What is the single biggest mistake firms make in season? A: Trying to deploy a new AI tool in March. Tooling decisions belong to the off-season. Once season starts, run the stack you have and capture lessons for next year.


If you want a season-ready AI stack scoped against your firm — your tax-prep system, your client mix, your head-count gap — reach out. We will pull last season's metrics, identify the highest-leverage tools, and have you ready by January. Or see the engagement on AI for accountants.

SOURCES

Cited and consulted.

  1. 01Journal of Accountancy — Tax Coveragejournalofaccountancy.com · accessed May 8, 2026
  2. 02Accounting Today — Tax Practice Coverageaccountingtoday.com · accessed May 8, 2026
  3. 03CPA Practice Advisor — Tax & Compliancecpapracticeadvisor.com · accessed May 8, 2026
  4. 04Karbon — Practice Management Librarykarbonhq.com · accessed May 8, 2026
View All Insights
NEXT STEP

Ready to ship the next outcome?

One Frequency Consulting brings 25+ years of technology leadership and military discipline to every engagement. First call is operator-grade scoping — sixty minutes, no charge.