Journal / Architecture
Why Tax Calculations Should Be Deterministic
AI can help build tax software. Explicit facts, rules, and reproducible calculations should produce the return.
In this essay
Ask a tax program to calculate the same supported return twice, using the same established facts and the same rules. The result should not depend on how the question was phrased, which explanation appeared earlier in a conversation, or whether a model found a different plausible answer this time.
That is the basic case for deterministic calculations. Given the same relevant inputs and rule version, the system should produce the same relevant result.
It sounds uncontroversial. But as AI becomes more capable, it is worth being precise about where that capability belongs. A model can help research a rule, propose an implementation, write a test, inspect a workflow, or challenge an explanation. None of those activities requires making the model the final runtime authority over a tax number.
PrepReturns was built with substantial AI assistance. Its tax calculations are explicit software. Those choices support each other: powerful development tools help construct and challenge a system whose output can be reproduced.
Reproducible is necessary, not sufficient
A deterministic program can be consistently wrong. It can implement the wrong rule, use an incorrect input, mishandle an edge case, or faithfully repeat a rounding defect. Reproducibility does not certify tax accuracy.
What it provides is a stable object to investigate. If an expected result and an actual result disagree, the developer and reviewer can work from the same inputs, identify the relevant rule, and reproduce the failure. A correction can then be tested against the case that exposed it and against cases that already worked.
Without that stability, diagnosis becomes entangled with variability. Was the difference caused by a fact, a rule, a prompt, a model update, or a different interpretation of the question? A professional calculation should not need that many possible explanations for a changed number.
Determinism is the beginning of an accountability structure. Evidence, rule research, independent expectations, regression tests, and human review have to supply the rest.
Facts, decisions, and calculations are different jobs
One source of architectural confusion is treating everything needed for a return as a calculation.
Reading a financial statement is not the same task as deciding whether its classification is appropriate. Establishing an asset’s history is not the same task as applying a supported depreciation calculation. Deciding how an item should be treated is not the same task as carrying that decision through a form.
These activities interact, but they deserve separate responsibilities. A source workflow proposes and preserves information. A review workflow establishes facts and records professional decisions. The engine applies supported rules to those established inputs. Workpapers and forms show the resulting relationships.
This separation does not eliminate judgment. It gives judgment a recognizable place. A preparer can change a supported election or resolve an uncertain fact, and the engine can consistently calculate the consequences.
“The software calculates it” should never conceal “the software guessed which facts and decisions you meant.”
Money deserves explicit handling
Even apparently simple arithmetic involves choices. A source may provide cents while a form presents whole dollars. An allocation may create a residual cent. Individually rounded components may not sum to the same display amount as a rounded aggregate.
Consider an illustrative allocation of $100.01 among three equal shares. Displaying three identical amounts is not enough to preserve the total to the cent. The system needs a defined rule for assigning the residual. It should use that rule consistently rather than depending on incidental processing order.
The same discipline applies when presenting a whole-dollar return alongside a detailed workpaper. The relationship between them should be explicit. A difference caused by display rounding should not be concealed with an unexplained plug merely to make the screen look tidy.
PrepReturns’ early tax-engine work addressed exact money handling, allocation, rounding, and the distinction between detailed and displayed amounts. This is unglamorous engineering, but it belongs near the foundation. A beautiful explanation cannot compensate for an unstable arithmetic policy.
The return is a network of consequences
A tax engine is more than a collection of independent form fields. One established fact can affect several outputs. A change to an asset, an owner, or a business classification can have consequences across the corporate return, supporting schedules, shareholder information, and workpapers.
That is why “the number on this line looks reasonable” is a weak acceptance standard. The relationships also need to hold.
A useful test might establish that a particular change affects the intended output while preserving unrelated results. Another might compare shareholder allocations with the corporate totals. Another might ensure that tax-basis work is not silently replaced by book-equity information simply because both are labeled with an equity-related word.
The engine should have explicit dependencies and clear domain distinctions. It should not require a reviewer to infer the entire computation from a sequence of screen updates.
For PrepReturns, this is part of treating the S corporation as a defined domain. The Form 1040 direction will carry forward those principles while using a separate model for households and individual tax questions.
Tests need expectations outside the implementation
It is easy to write a test that repeats the program’s logic and confirms that it produces the same answer as itself. That can catch accidental changes, but it is a weak challenge to the underlying reasoning.
Stronger tests begin with independently established expectations. For a synthetic scenario, the author can define the facts, derive expected relationships, and identify boundaries before comparing the engine’s result. A fresh reviewer can question the derivation instead of simply accepting the existing output as the answer key.
This does not mean building a second complete tax engine to verify the first. It means selecting meaningful independent checks: literal expected values, cross-form relationships, boundary cases, known failures, and changes whose consequences can be reasoned about separately.
A good regression fixture is a durable question: “Does this still behave as intended?” As scope expands, the collection of questions becomes part of the product’s engineering memory.
Synthetic examples are particularly useful because their facts can be complete and intentionally designed. They are evidence about those cases and checks, not a count of client returns or a guarantee about every possible return.
A missing fact is not an invitation to improvise
Deterministic software needs a policy for incomplete inputs. The least useful policy is to invent a convenient default and continue as if the return were complete.
If the system requires an opening amount and that amount is unknown, the uncertainty is part of the case. A zero is a substantive value. It should not be manufactured by an empty text box or an unavailable source.
Similarly, a capability boundary should be represented as a boundary. A conventional supported transaction and a superficially similar unsupported transaction may require different rules. Producing a plausible-looking form in both cases does not make the product more capable.
Explicit unknowns and supported-scope checks complement deterministic calculation. They answer the question that comes before arithmetic: is the engine entitled to calculate this result from these facts?
This is why the S-Corp Engine’s boundaries are part of its product definition. A smaller, explained calculation universe can be expanded and reviewed deliberately.
Saved output needs the same discipline
Reproducibility also concerns which calculation a user is seeing. A return package saved yesterday should not silently become today’s package after a fact changes.
Suppose a preparer changes an owner address or corrects an amount after producing a return. The old PDF may remain useful as history, but it should not look like the current output. Diagnostics should not describe one preparation while Forms displays another.
PrepReturns’ saved preparations and current-versus-historical distinction address that relationship. Updating the return creates a new prepared package rather than rewriting the meaning of the old one. Forms, workpapers, and review results need to refer to the same saved preparation.
That is a broader form of determinism: the system should preserve the identity of the work, not merely repeat its arithmetic. A reviewer should be able to say which facts and preparation produced the document in front of them.
The visible form is another engineering problem
A correct calculation does not guarantee a correct-looking PDF. Field mapping, page selection, text appearance, and package assembly create another set of possible failures.
During Alpha 14 development, direct inspection exposed a particularly useful example: official-form mappings could exist while text failed to appear visibly after PDF page selection. The repair concerned how the PDF was serialized and flattened before selecting pages, rather than a change to tax arithmetic.
This is why output review cannot stop at checking that a field contains a value. Someone needs to inspect what the preparer actually sees. A blank visible field is a product defect even when a lower-level test says the data is present.
The lesson is not that deterministic calculations are insufficiently ambitious. It is that professional output requires several layers of correctness, each with its own verification. Computation, saved-state identity, rendering, and human readability all matter.
Use AI where its strengths help
Frontier models are useful development collaborators precisely because they can explore alternatives, explain unfamiliar code, propose tests, and challenge assumptions. Those activities benefit from flexible reasoning.
They also require supervision. A model can misunderstand a rule, write a test around its own mistake, or produce a persuasive explanation of incorrect code. Fresh review and visible output checks help challenge those failures; the human professional still sets the scope and acceptance standard.
That is distinct from runtime source extraction. Using AI to write an importer does not mean that importer uses AI when it runs. The current project’s supported source paths should be described on their own terms, without borrowing capabilities from the tools used during development.
The durable boundary is straightforward: development assistance can be expansive, while established facts and explicit rules remain the basis of the calculation. The technology page shows how these responsibilities fit together.
Make the result worth trusting
Trust is often discussed as if it were a feeling generated by a polished interface. In professional software, it should be supported by things a reviewer can examine.
Can the result be reproduced? Can the relevant facts be found? Can a disputed treatment be changed deliberately? Is the rule behavior testable? Does a correction preserve history? Does the visible output match the saved preparation? Does the software admit when required information or support is missing?
Deterministic tax calculations make those questions easier to answer. They do not remove tax judgment or promise infallibility. They give the professional a stable system to inspect, challenge, and improve.
AI helped make PrepReturns possible. Explicit calculation logic makes its results something the preparer can reason about. The ambition is to use both strengths well, with human authority visible throughout the development process and the preparation workflow.