
Two practitioners ask the same AI tool the same tax question and get answers of wildly different quality. The tool didn't change between them. The prompt did.
That is the whole game. Prompting AI for tax research is a discipline, and the difference between a throwaway answer and a defensible one lives almost entirely in how you frame the request. Most of us learned to search: type keywords, scan results, synthesize. That reflex actively works against you here. These tools reason over what you give them, so a thin prompt produces thin reasoning, and a leading prompt produces exactly the answer you nudged it toward, right or wrong.
This is a prompt engineering playbook. We will start with why the mechanics differ from search, break down the anatomy of a strong prompt component by component, and then run weak-versus-best comparisons across the tasks you actually do: issue spotting, multistate questions, memos, and stress-testing a position before it goes out the door.
Start with an accurate mental model of the thing you are talking to. A search engine matches your keywords against an index and hands back documents. An AI research tool reasons over the facts and instructions in your prompt and hands back an analysis. Those are different machines, and they reward opposite habits.
A search query gets shorter and more keyword-dense as you refine it. A good research prompt gets longer and more specific. You are not scanning an index. You are briefing a junior researcher who is fast, well read, literal, and entirely dependent on the facts you choose to include. Anything you leave out, it will either ignore or quietly invent.
That reframing produces the core rule: a prompt is not a query, it is a fact pattern plus an instruction set. Everything below is a way of making the fact pattern complete and the instructions explicit.
Here is the pattern across common tasks. Read the middle column first: the reason a weak prompt fails is almost never that it is too short. It is that it omits the facts, the year, the jurisdiction, or the standard of proof the answer depends on.
Every rewrite does the same handful of things. Let us break down what those things are.
Five components turn a question into a research instruction. Miss one and you have handed the tool a decision to make on your behalf.
Brief the tool the way you would brief a senior before they open the file: filing status, entity type, dollar amounts, dates, relationships, residency. The single biggest quality lever is facts, not phrasing. When you leave a fact out, the model fills the gap with an assumption you never see and never approved, and the answer reads just as confidently either way.
One caution that belongs here and nowhere later: loading facts means loading client facts, which raises a confidentiality duty before you paste anything. More on that below.
Tell the tool who it is being. "You are a senior tax researcher preparing a position for partner review" produces more rigorous, more hedged, more citation-driven output than the same question asked cold. Role framing sets the standard of proof. Compare:
The second version gets a scoped, current, audience-aware answer instead of a textbook dump.
Say what level of authority you will accept. Statute, regulation, revenue ruling, IRS publication, a named state code section. Asking for "the controlling authority" and telling it to distinguish binding law from sub-regulatory guidance pushes the tool toward primary sources and away from paraphrasing whatever secondary commentary it has seen most often.
A file memo, a bulleted issue list, and a client email are three different deliverables. Name the one you want, its length, and its reader. "Explain this" gets you shapeless prose. "Draft a two-paragraph plain-English explanation for the client, no citations in the body" gets you something you can almost send.
Ask for the chain from fact to authority to conclusion, not just the conclusion. When the reasoning is visible you can audit it, and you can see the exact step where a wrong turn happened. An answer with no visible path is an answer you cannot review, which means it is an answer you cannot use.
The components above build one strong prompt. These techniques work across the whole research session.
If you want output in a specific shape, show the tool one example of that shape before asking for the real thing. Paste a short sample memo, then say "produce the analysis in this exact format." Few-shot prompting is the fastest way to lock structure without describing it in tedious detail.
Treat the first response as a draft, not a deliverable. The highest-value prompting often happens on turn two: "You assumed the activity rises to a trade or business. Redo the analysis assuming it does not." Each follow-up narrows the answer toward your actual facts. A conversation beats a single perfect prompt almost every time.
The most underused prompt in tax research is the adversarial one:
"Argue the opposite of the conclusion you just reached. What fact or authority would an IRS examiner use to challenge this position, and where is it weakest?"
Turning the tool against its own answer surfaces the soft spots before a reviewer or an examiner does. It is the closest thing to a second set of eyes late in busy season, and it costs you one line.
For a multi-issue return, do not ask one sprawling question. Split it: spot the issues first, then take them one at a time, then ask for the memo that ties them together. Decomposition keeps the reasoning clean and makes each step reviewable on its own.
The failure modes are consistent and every one is avoidable.
Tools like Marble's Intelligence agent are built for exactly this loop: ask a federal or state tax question in plain English, get a citation-backed answer that links directly to the controlling IRC section or IRS guidance, then turn it into a client-ready memo you can review and sign. The prompting discipline in this post is what makes any research tool worth trusting. Better inputs, more defensible output. Join the Marble waitlist.
The completeness of the fact pattern. Filing status, entity type, dollar figures, dates, and the tax year do more for answer quality than any phrasing trick. The tool reasons from what you give it, so a missing fact becomes a silent assumption.
A weak prompt is a keyword-style query that omits the facts, year, or jurisdiction the answer depends on. A best prompt is a full fact pattern plus explicit instructions on authority, format, and reasoning. The rewrite is rarely about being longer for its own sake. It is about removing every decision the tool would otherwise make for you.
Yes, without exception. A citation shows where the tool says the authority lives, not that the authority supports the conclusion. Open the primary source, read the relevant passage, and confirm it says what the answer claims. That step is faster than untangling a wrong position after it has gone out.
Only after confirming how the tool stores, retains, and uses the data. Disclosing taxpayer information can implicate Section 7216 and your firm's obligations under the FTC Safeguards Rule. Vet the data handling first, and strip identifying details when the underlying facts are all the research requires.