Journal / Development
What Can a Small Tax Firm Build Now?
PrepReturns explores how frontier AI changes the feasibility of building professional software while leaving the hard work intact.
In this essay
A small tax practice frustrated with its software has traditionally had a familiar set of options: change a process, build a spreadsheet, request a feature, investigate another vendor, or accept the friction.
Building an end-to-end preparation system was a different category of ambition. The practice might understand the problem exceptionally well and still lack a practical way to coordinate all the software work around it.
Frontier AI-assisted development changes that calculation. It makes research, implementation, testing, review, and iteration more accessible to a small organization with a concrete problem and the judgment to keep pursuing it.
PrepReturns is an experiment in that possibility. It grew out of Trio Tax, a working tax practice. The first engine focuses on conventional S corporations, with Form 1040 next. The interesting result is not that a model can generate code. It is that a practice can now attempt to assemble the whole preparation path, from source evidence to actual forms and a saved return package.
Domain knowledge becomes a development resource
Practitioners know things that are difficult to discover from an abstract product brief. They know where a handoff is awkward, which screen hides context, which missing fact matters, and why a superficially convenient shortcut would make review worse.
That knowledge does not automatically translate into good software. But better implementation tools give it a more direct route into the product.
A request such as “make source review easier” is broad. A practitioner can make it concrete: show the accounts contributing to a payroll disagreement, distinguish an unread amount from a known zero, preserve a correction reason, and return to the unresolved preparation task after the issue is addressed.
Each refinement reduces ambiguity for the people and models doing the implementation. The practice supplies the meaning of the work; development tools help turn that meaning into a system that can be exercised and challenged.
This is one reason the origin matters. PrepReturns began with preparation problems, not with a software category looking for a market.
A narrow domain can still demand a serious system
Choosing S corporations first bounded the project. It did not make the work small.
A recognizable operating business can involve multiple owners, payroll, assets, distributions, shareholder debt, retirement contributions, inventory, rentals, or investment sales. Supporting a deliberate subset of those patterns still requires careful distinctions between books, tax facts, corporate reporting, and owner information.
The “straightforward 90%” describes a design ambition. It is not a measured coverage statistic or a claim that conventional returns require little judgment. The aim is to make a defined universe coherent instead of starting with every unusual case.
That boundary gives development a useful discipline. A proposed capability can be evaluated as a concrete expansion: which facts are required, which outputs change, which cases remain unsupported, and which existing behavior must survive?
Scope becomes a way to make progress responsibly. It should be clear enough that the software can recognize when the case exceeds it.
The first result had to be a calculation system
Early PrepReturns work concentrated on a deterministic tax core. Before a polished workspace could be convincing, the calculations needed explicit money handling, defined relationships, supported boundaries, and independently derived synthetic expectations.
This matters because a prototype can look complete long before it has a dependable domain model. Forms and input screens are easy to recognize. The distinctions behind them are less visible: book equity versus stock basis, a known zero versus missing history, an ordinary item versus a separately stated one.
The early engineering question was whether the tax logic could stand on its own and be challenged. A professional interface built on unstable meanings would only make the instability easier to operate.
The development history follows that foundation into later generations. The progression is cumulative: once a calculation exists, every new workflow must preserve the facts and rules on which it depends.
A calculator is only the beginning
The next challenge was to connect the engine with the information a preparer receives. That introduced structured intake, supported source parsing, financial-statement mapping, prior-return handling, and review.
Then the project needed a professional workspace around those activities: case storage, saved state, history, recoverability, and a usable browser interface. Guided Preparation added another question: how can the system ask for the next useful fact without making the preparer navigate the implementation’s internal structure?
Source corrections made the problem deeper still. A number read from a document is a proposal. A preparer may need to correct the reading, add a missed supported item, exclude an extra software reading, or classify the information differently. Those actions have different meanings and should leave different evidence.
This is where the phrase “AI built an app” becomes too small to describe the work. The work is deciding what the app must mean when a user does something consequential.
The agents need a job description
Coordinated models can help with several parts of that work. One assignment might concern a bounded capability. Another might examine a research question. Another might review output or challenge an implementation with fresh context.
The value comes partly from separation. A reviewer who starts with an expected behavior can question an assumption the implementing worker has become attached to. Parallel workers can examine different areas while a lead retains responsibility for shared contracts and integration.
But dividing work is itself work. Someone has to define boundaries, resolve overlaps, keep shared representations consistent, and decide which findings require changes. A collection of individually plausible contributions does not automatically make a coherent system.
The Alpha records describe coordinated implementation and review roles under human direction. Models were given problems to investigate and work to challenge, while the project retained a human definition of what the software was supposed to do. The useful lesson is the working method: bounded tasks, explicit expectations, integrated checks, and human authority.
A form can contain data and still look blank
One Alpha 14 defect explains the limits of code-level confidence particularly well. New official-form output could have the intended field mappings while visible text disappeared during PDF page selection.
The failure concerned the handling of PDF appearance data. The remedy involved serializing and flattening the filled PDF before selecting pages, followed by renewed inspection of actual output.
The important point is what exposed the defect. A check that a field existed or held a value was not enough. The form had to be rendered and looked at as a human would see it.
That lesson applies well beyond PDFs. A button can exist without leading to the right task. A table can contain correct values while hiding its headings at a common viewport. A successful network request can still leave a user unsure whether their work was saved.
The professional’s visible experience is part of the implementation. Browser and output inspection are a way of testing the product’s meaning, not just its appearance.
Integrity has a performance cost
Another Alpha 14 problem arose from a feature worth preserving: immutable history and verification. The richer synthetic Cedarline case accumulated enough history that repeatedly verifying it became expensive. Selected exports were redoing work, and browser operations could time out.
The easy response would have been to remove checks or prune inconvenient history. That would have changed the product’s promises. The engineering task was to reduce repeated work while preserving the verification boundaries.
The resulting work used bounded reuse of verified structures and explicit checks on saved-file membership, paths, and hashes. It required reasoning about which proof remained valid, which dependencies still had to be examined, and when a selected output could safely be assembled.
This is a useful example of the project’s cumulative difficulty. Once evidence integrity becomes a product feature, performance changes must respect it. An optimization is not successful merely because a stopwatch improves; it must preserve the meaning of the operation.
The historical measurements belong to their particular synthetic case and environment. They are engineering evidence, not universal current performance benchmarks.
Tests need several kinds of questions
As the engine expanded, the testing obligation expanded with it. Pure calculation checks address a different question from source-workflow checks. Saved-history checks address a different question from whether a shareholder PDF contains the right pages.
A useful test strategy needs those layers. Literal expected amounts challenge arithmetic. Integrated scenarios challenge relationships. Deliberate omissions and conflicts challenge intake. State checks challenge stale and historical output. Browser inspection challenges the workflow. Actual PDF inspection challenges the rendered package.
The synthetic businesses provide a ladder of increasing complexity. Juniper demonstrates the core path. Alder adds multiple-owner and basis interactions. Cedarline adds richer operating-business features, including inventory, rental, vehicle, investment, and amortization work within the supported boundaries.
These examples let reviewers know which facts were supplied and which results were expected. They do not represent client-return usage counts. Passing checks over them does not prove every tax rule or every real-world source layout.
The point of the evidence is to make the development work assessable, with its scope intact.
Human review does not disappear at the end
Models can find mistakes made by other models. They can inspect a browser, question an answer key, and propose a repair. That is useful, but it is not external professional certification or a replacement for the person responsible for the practice.
The human owner defines the problem, decides the intended scope, and evaluates whether the workflow supports actual preparation. A technically consistent interaction can still be professionally awkward. A supported calculation can still require better questions before the preparer should rely on it.
Alpha 14 is the human-testing generation and broadens conventional operating-business coverage. That description preserves both the progress and the continuing role of practitioner review. It should not be inflated into an undocumented claim that every acceptance step has been completed.
The software is meant to serve professional judgment. Its development process should make that relationship visible too.
The next domain should keep the lessons
Form 1040 comes next. That means a new domain, not simply more forms attached to an S-corp object.
Households, people, dependents, wages, investments, deductions, credits, and individual limitations have their own relationships. The useful inheritance is architectural: source provenance, explicit unknowns, deterministic calculations, direct workpapers and forms, saved history, and professional control.
The business-to-owner connection is a good example. Structured shareholder information should eventually be able to move across a defined boundary. The individual engine must still own the individual questions and consequences. Sharing information should not require pretending that corporate and household preparation are the same problem.
The 1040 direction is therefore a complete next-stage idea, not a claim that an individual engine already prepares returns.
The opportunity is greater agency
PrepReturns does not establish that every tax practice should build a tax engine. Different firms have different needs, skills, priorities, and tolerances for ongoing software responsibility.
It does suggest that the available choices have changed. A practice can now attempt more than a spreadsheet workaround. It can explore a workflow, construct a bounded system, and develop evidence about whether its ideas work.
The project remains private today, with an open-source future as the intended direction. That future could let others inspect, integrate, modify, or contribute under the terms of an eventual release. Trio Ledger, the sibling small-business ledger product, reflects the same interest in building from practical needs outward; no live integration is implied by the relationship.
What can a small tax firm build now? PrepReturns is our concrete attempt to answer. The answer includes working software, difficult defects, expanding tests, and continuing professional review. It also includes an invitation to compare notes about what becomes possible when the people closest to the work can participate more directly in building their tools.