Journal / Professional tools
Let Me Test-Drive Your Tax Software
A better way to evaluate professional software: prepare, correct, inspect, and ask technical questions in a synthetic case.
In this essay
A sales demo isn’t a test drive.
A demonstration can be useful. It introduces a product’s vocabulary, shows the intended workflow, and gives an evaluator a map. A thoughtful presenter can explain why the software is organized the way it is and point out capabilities someone might otherwise miss.
But watching someone drive a familiar route does not tell me how the controls feel when I use them. Professional tax software is something a preparer has to think inside. To evaluate it, I want to do the work: establish facts, change my mind, inspect a calculation, resolve a diagnostic, and produce output I can read.
That is not an unreasonable request for mission-critical software. It is a different evaluation method, with different evidence. I think the industry would benefit from making that method more normal, and it is a standard I want PrepReturns to work toward as well.
A guided route answers only some questions
A presenter knows where to click, which fields are important, and which path makes a feature intelligible. That knowledge is valuable. It is also precisely what a new user does not yet have.
When I watch a demonstration, I can assess broad organization and visible capabilities. I cannot reliably assess whether I would have found the next action without help, understood a warning, or noticed that a result belonged to an earlier preparation. The demonstration substitutes the presenter’s experience for part of my own.
Even a perfectly honest presentation has this limitation. There is only so much time, and no meeting can cover every working style or unfamiliar case. The issue is not whether the presenter is concealing something. The issue is whether the evaluation supplies the evidence needed for a decision.
For a tax practice, the cost of choosing software extends beyond its price. There is learning, process change, conversion work, review, and the disruption of finding an awkward fit after commitment. Hands-on evaluation cannot remove all of that uncertainty, but it can move some discoveries earlier.
I would like the sales conversation and the test drive to complement one another. Show me the intended route. Then let me explore it myself.
The case should be synthetic and believable
A useful evaluation does not require real taxpayer information. A carefully constructed synthetic business can supply books, prior-year context, owner information, assets, and supporting records without asking anyone to expose a client.
The quality of that case matters. If every fact arrives perfectly classified and every decision has already been made, the evaluator mostly tests data entry. If the packet is chaotic in ways unrelated to the software’s intended scope, the evaluation becomes an exercise in guessing what the author meant.
The best case gives the preparer enough information to establish a result while leaving meaningful work to perform. Include a source value that needs review, a supported correction, an unresolved professional decision, and a reason to inspect output. Make the business recognizable. Explain what information is fictional and what behavior the case is intended to exercise.
It is useful to have a ladder of cases. A simpler case reveals first-use clarity. A multiple-owner case reveals allocation and output organization. A richer operating-business case reveals how the system handles interacting activities. The evaluator can choose the level that matters without starting in the deepest part of the product.
PrepReturns uses Juniper, Alder, and Cedarline as synthetic examples along that progression. The public website shows selected screenshots and output from those examples. A hosted test environment is a future direction, not a destination that this site pretends already exists.
Let me make a correction
The most revealing action in a software evaluation may be changing something that looked finished.
Start with a proposed fact from a supported source. Review it, establish a different value with a reason, and prepare the return. Then revisit the source. Can I still see what it originally said? Can I distinguish my correction from the parser’s proposal? Does the calculation use the established fact? Can another reviewer understand the difference?
Next, change that established fact after output exists. Does the software clearly distinguish the current inputs from the historical package? Does it ask for a new preparation where appropriate? Can I open the old package as history without mistaking it for the updated return?
These actions reveal a system’s model of authority and time. They are difficult to judge from a static feature list. A product can truthfully advertise document import, overrides, and PDF output while the relationship among those features remains awkward.
I want to explore that relationship because preparation is iterative. Review produces questions. New information arrives. A source can be wrong, and a preparer can be wrong. Good software should help recover while preserving the meaning of earlier work.
The evidence model in PrepReturns grew around exactly these questions. It is one of the areas I most want practitioners to challenge.
Let me follow a number
An evaluation should include a number whose explanation is worth inspecting. Where did it come from? Which facts affected it? Which treatment decision mattered? How does it appear in a workpaper and on a form?
The point is not to demand that every screen expose implementation detail. Most preparers do not need a source-code listing beside a form line. They do need an explanation suited to the professional question they are asking.
For example, an owner-specific result should make it possible to understand the relevant ownership information and supporting computation. A return total should connect to its underlying categories. A diagnostic should tell the preparer what relationship failed or needs attention, rather than merely reporting that the software is unhappy.
The evaluator should also inspect an actual exported package. Is the text visible? Are the pages legible? Are supporting statements included where expected? Is owner information separated appropriately? Does the output correspond to the preparation being reviewed?
Those may sound like basic requirements. They are still worth testing. In PrepReturns’ own development, direct PDF inspection found a rendering defect that correct field mappings had not prevented. Looking at the artifact mattered. The development record explains why output inspection became an explicit part of the loop.
Give the technical questions a real destination
Salespeople should not have to improvise answers to architecture questions they were not hired to resolve. A good evaluation process gives those questions a route to someone who understands the product technically.
I might want to know what happens to stored history after an update, whether an export includes provenance, how an integration distinguishes a proposed value from an approved fact, or what prevents a stale package from being presented as current. Those questions are not objections to overcome. They are part of understanding fit.
A useful answer can be “the current product does not support that,” followed by a clear explanation of the available path. It can also be a written follow-up from an engineer. What matters is that the answer is specific enough to inform a decision and does not change depending on who is asked.
Technical documentation helps, especially when it describes behavior and boundaries rather than merely naming features. A short example of how a system handles corrected evidence can be more useful than a slide promising automation. A concrete explanation of exports can be more useful than the word “integrations.”
The purpose is not to turn every purchase discussion into an engineering review. Different firms need different depths. The deeper path should exist for people whose decision depends on it.
A fair test includes the intended scope
A test drive should not be a trap. If a product is designed for conventional S corporations, evaluating it solely through an unusual transaction outside that scope does not tell us much about whether it succeeds at its purpose.
At the same time, scope should be visible before the evaluator reaches the unsupported situation. The vendor should explain the intended universe and the system should communicate boundaries clearly. “You should have known that was unsupported” is not a satisfactory user experience when the interface appeared to accept it normally.
A fair evaluation therefore has two parts: exercise a representative supported workflow, then explore how the system behaves at its edge. Does it identify missing information? Does it distinguish unsupported circumstances from incomplete input? Does it provide an intelligible stopping point?
This is also a way to evaluate professional control. Can the preparer inspect the forms directly? Can they navigate back to source evidence? Can they understand why a choice is unavailable? Guidance is helpful when it preserves orientation. It becomes frustrating when it requires blind trust.
Keep the evaluation reproducible
For the experience to be useful, the evaluator should be able to start again. A fresh synthetic case lets someone distinguish their own edits from the initial state, repeat a correction, or compare two approaches.
A small answer guide can explain the intended results and important review questions. It should not silently pre-resolve every professional decision. The evaluator needs to know which facts are supplied and which conclusions they are expected to establish.
There is value in letting people leave the happy path and recover. Open a different screen. Make a reversible mistake. Try a supported edit before finishing the interview. Inspect the print selection. A product used under deadline pressure will encounter these patterns, whether or not a demonstration anticipates them.
Reproducibility also makes feedback better. “I couldn’t understand this” becomes more actionable when both parties can reset the same case and follow the same steps. It gives designers and engineers a shared artifact rather than a vague account of a confusing moment.
This is a product obligation too
It would be easy to direct all of these expectations at other vendors. Building PrepReturns makes them obligations for me as well.
The public site needs to show real software, identify its synthetic examples, describe its boundaries, and provide enough architecture and development detail for a serious reader. It should not send someone to an invented demo or imply that screenshots are a substitute for hands-on qualification.
The S-Corp Engine pages are a way to inspect the current work. The roadmap describes a synthetic hosted showroom as future work. Until that experience exists, the honest path is to explain what can be evaluated publicly and invite a conversation about the rest.
I want professional software to be easier to understand before a firm commits to it. A good demonstration opens that conversation. A test drive makes it concrete. Let the practitioner prepare, correct, inspect, and ask a difficult question. The resulting feedback may be more useful to the product than the smoothest sales meeting ever could be.