Your Client's AI Project May Be Creating an R&D Tax Credit Opportunity
A practical guide for AI consultants, agencies, and fractional CTOs: technical experimentation, contract risk, and the evidence clients should preserve before year-end.
AI consultants are moving traditional businesses into territory those businesses may never have associated with research and development. A distributor that once purchased software is now building a custom document-processing system. An insurance agency is testing ways to extract and validate information from carrier documents. A professional-services firm is developing its own method for matching information across emails, meeting notes, phone records, and document systems.
These companies may not think of themselves as technology companies. Yet some of the work they are undertaking may create a legitimate federal R&D tax credit opportunity.
The AI consultant is often the first adviser in a position to recognize it. You understand what the client is trying to build, where the available tools fall short, which technical questions remain unresolved, and how the team plans to test possible solutions. Raising the issue early can help the client evaluate the opportunity while the project, contract, and supporting evidence are still taking shape.
The opportunity often begins when buying turns into building
Purchasing an AI license does not, by itself, create an R&D credit. Neither does deploying a standard chatbot or configuring a commercial platform to perform functions it already supports.
The analysis changes when the available product cannot meet the client's requirements and the client or its consultant begins developing something new around it. That might include a custom data pipeline, retrieval architecture, validation layer, evaluation harness, integration, agent workflow, or performance improvement.
The strongest opportunities generally share a recognizable development story:
- The business needs a new or improved function, level of performance, reliability, or quality.
- The team does not know at the outset whether the result is technically achievable, which method will work, or what the appropriate design should be.
- The developers identify alternatives and evaluate them using testing, modeling, simulation, or systematic trial and error.
- The results, including failures, change the technical design.
A practical question for an AI consultant to ask is:
What did the client need the system to do, what technical answer was unknown at the beginning, which alternatives were evaluated, and what results changed the design?
If the project has a substantive answer, it deserves a closer look.
Five AI development patterns worth recognizing
1.Building a custom layer around a purchased platform
Suppose a company buys a document-processing platform that achieves only 65% accuracy on its records. The developers create custom preprocessing, classification, validation, and exception-handling layers. They build an evaluation dataset, compare five approaches, and eventually reach 95% accuracy.
The potentially qualifying work is not the platform purchase. It is the custom technical development and evaluation undertaken to resolve the accuracy problem. License costs, user training, and routine configuration should be separated from that work.
2.Resolving identities across inconsistent data sources
A business wants to connect emails, meeting transcripts, phone notes, and stored documents to the correct customer. No shared identifier exists across every system. The team tests combinations of names, addresses, email addresses, telephone numbers, and source-specific weighting methods. It measures false matches and missed matches, then changes the design for different ingestion sources.
This is more than ordinary data cleanup when the team is resolving genuine technical uncertainty through repeatable testing. The qualifying portion may include the matching architecture, source-specific experiments, automated ingestion methods, and evaluation process, not routine manual corrections after the system is operating.
3.Developing a RAG system with measurable requirements
A client needs an internal knowledge system to return accurate answers within a strict response-time limit. The team does not know whether vector retrieval, a knowledge graph, hybrid search, or reranking will satisfy both requirements. It creates a representative question set, measures retrieval quality and latency, and compares architectures.
The fact that the project uses retrieval-augmented generation does not prove eligibility. The meaningful facts are the technical uncertainty, defined alternatives, evaluation method, and results that guided the architecture. If the system primarily supports general administrative functions, the internal-use-software rules may add another layer of analysis; "used internally" is not an automatic yes or no.
4.Improving a vertical AI application
An AI builder develops an industry-specific application that must hold hallucinations below a defined threshold. Its team evaluates multiple models, prompt structures, retrieval methods, context strategies, and guardrails against a labeled dataset.
Prompt work is not automatically qualified or disqualified. Casual prompting until an answer sounds better is very different from a structured computer-science experiment with a technical target, controlled alternatives, repeatable evaluation, and documented results. The development process, not the word "prompt," drives the analysis.
5.Developing for a client under a services agreement
An AI agency agrees to build a system for a client. Which party has the stronger potential credit position can depend on several facts: who absorbs the cost if the technical research fails, what rights each party retains, and where the developers work. The parties must also separate qualified research from the engagement's other services.
This is why an agreement should be reviewed before development begins. "Fixed fee," "time and materials," and "the client owns the IP" are useful facts, but none answers the question alone.
Use the Need–Unknown–Alternatives–Evidence screen to determine whether the project deserves an early R&D credit conversation, before the contract, development process, and supporting evidence are fixed.
Run the five-question self-screenA practical view of the qualification rules
The federal framework is sometimes called the four-part test. In practical terms, the work must relate to a new or improved business component and pursue better function, performance, reliability, or quality. It must also address technical uncertainty through a process that evaluates alternatives using computer science, engineering, or another qualifying hard science.
Qualification attaches to activities and business components, not industries or company descriptions. A software company does not qualify merely because it employs developers. A construction company, medical practice, accounting firm, or distributor is not excluded merely because technology is not its primary product.
The development also does not need to be new to the world, patented, successful, or performed in a laboratory. A failed project can still contain qualified research when the failure arose from a genuine experimental process. Conversely, a successful and expensive project may contain little qualified research if the team simply implemented a known solution.
When the client hires the AI consultant
A client may generally treat 65% of the qualifying portion of eligible contractor payments as contract research expense when several contract-research conditions align. The agreement must exist before the research, and the work must be performed on the client's behalf. The client must have rights to the research results and bear the expense even if the research is unsuccessful. The underlying activities must satisfy the qualification rules, and the claimed federal research must be performed in the United States, Puerto Rico, or another U.S. possession.
The 65% rule does not apply automatically to the entire invoice. If an engagement includes experimental development, ordinary configuration, deployment, training, and maintenance, the qualifying development portion must first be identified.
For example, assume $160,000 of a $200,000 engagement relates to potentially qualified development. The remaining $40,000 covers training and routine implementation. The starting contract-research amount would generally be 65% of $160,000, or $104,000, not 65% of the full invoice.
That $104,000 is not the credit. It is a qualified research expense that enters the broader credit calculation, which depends on the taxpayer's history and computation method; as a rough working range, the federal credit often lands around 6% to 10% of qualified expenses. On contractor spending alone, that example supports something in the neighborhood of $6,000 to $10,000, which is real money but may not justify a full readiness engagement by itself. The economics usually become compelling when internal wages join the picture.
An illustrative build, end to end
Consider a hypothetical mid-sized business, not a client, with three US developers spending most of a year on a genuinely experimental AI system: roughly $450,000 of qualifying wages, the $104,000 of qualifying contractor expense above, and $30,000 of cloud compute used in development. That is in the range of $580,000 of qualified research expenses, and at the 6% to 10% working range, a federal credit somewhere between $35,000 and $58,000, potentially recurring in each year the qualifying development continues. For a young company that elects the payroll-tax offset on an original, timely filed return, that figure can arrive as near-term cash rather than a carryforward.
Every number above is illustrative, and an honest evaluation cuts both ways: we recommend proceeding past an assessment only when the conservative benefit clearly exceeds the total cost of the work, including the independent specialist's study. Sometimes the numbers say no. That discipline is what makes the yes worth something.
The AI consulting firm may have its own opportunity
The client is not always the party with the stronger position. Under the funded-research rules, an AI agency, engineering firm, or development consultancy may potentially claim its own qualifying expenditures. The firm generally needs to bear the financial risk of research failure and retain substantial rights in the results.
Consider an agency that earns most of its fee only if a system meets a defined technical acceptance standard. The agency pays its U.S. developers during unsuccessful iterations, must perform rework at its own expense, and retains the right to reuse its underlying framework and generalized technical components. Those facts can point toward the agency bearing the research risk and retaining meaningful rights.
The complete agreement matters. Payment milestones, acceptance provisions, rework obligations, refunds, termination payments, warranties, rights to newly created and pre-existing technology, reuse rights, and developer locations can all affect the analysis. It is also possible to structure risk and rights so poorly that neither party has a clean credit position.
An early conversation does not mean rewriting an agreement solely to chase a tax result. It means understanding the commercial arrangement the parties actually want and recognizing its tax consequences before the facts are fixed.
Why the conversation should happen during development
The R&D credit is much harder to support when everyone waits until after year-end and tries to reconstruct what occurred.
During the project, the evidence usually already exists: technical plans, architecture decisions, evaluation datasets, test results, failed approaches, development tickets, employee allocations, contractor invoices, agreements, and work-location records. A practical readiness process connects that ordinary project evidence to the relevant business component and technical activities.
Documentation does not turn routine implementation into qualified research. It preserves the facts supporting work that may already qualify, and helps separate it from deployment, training, maintenance, and other ordinary activities.
RHW is applying this same readiness process to its own 2026 entity-resolution development, including the technical alternatives evaluated, the results that changed the design, and the evidence preserved during development.
Why a CPA who understands AI development is an advantage
AI projects do not arrive in the vocabulary of the tax law. They arrive as RAG pipelines, model evaluations, entity-resolution problems, validation layers, confidence thresholds, latency constraints, and iterative architecture decisions.
Working with a CPA who understands both the tax framework and how AI systems are built reduces the translation gap. The consultant can explain the development in its natural technical language. The CPA can identify the corresponding questions about business components, uncertainty, experimentation, internal use, qualified expenditures, contractor risk, research rights, and evidence.
That combination also supports better judgment. The answer is often not that an entire company or project qualifies. It may be that one development phase qualifies, another requires additional facts, and routine implementation should be excluded. Recognizing those distinctions early produces a more credible analysis and a cleaner handoff to the specialist responsible for the formal credit study and filing conclusions.
RHW works with AI consultants to help their clients identify legitimate R&D credit opportunities early, evaluate the relevant facts, and preserve the evidence needed to support a potential credit.