Tracking R&D Capitalization for AI Engineering Work That Doesn't Look Like Software
R&D capitalization tracking was built around a mental model of engineers writing code in a repository. AI engineering work often looks different: data curation, training runs, evaluation loops, and prompt iteration, much of which doesn't produce a commit an accountant can point to as evidence of progress.
That doesn't change what's capitalizable, it changes how you need to track it, since the usual proxy of "commits per sprint" doesn't capture what an AI team is actually spending its time and budget on.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Why AI work breaks the usual tracking habits
Traditional software capitalization tracking often leans on time tracking tied to tickets or sprints in a project management tool, on the assumption that a ticket roughly maps to a deliverable. AI development work includes plenty of time that doesn't map cleanly to a ticket: running an experiment that fails, tuning a training configuration, or waiting on a long training job to complete while reviewing results. None of that is wasted time from an accounting standpoint, but it needs its own tracking category rather than being lumped in as overhead.
Tracking headcount time against project stage
Have engineers log time against project and stage, not just project, using whatever project management tool you already run. The category that matters most is distinguishing exploratory work, deciding whether an approach will work at all, from committed build work, executing on an approach you've already decided to pursue. That's the same line ASC 350-40 has traditionally drawn between preliminary and application development stages (ASU 2025-06 replaces those stages with authorization and probable-completion criteria for annual periods beginning after December 15, 2027), so getting engineers into the habit of tagging work by project phase helps both operational visibility and capitalization support under either model.
Tracking compute cost as its own capitalizable input
Training and evaluation compute is often the largest single cost in an AI capitalization project, and it needs the same stage tagging as headcount time. Tag training jobs at the point of submission with the project and stage they belong to, so the compute bill can be split the same way the payroll cost is, rather than left as an undifferentiated infrastructure expense that never gets matched to the project it supported.
Rolling it up into total cost of ownership
The number a CFO actually wants isn't payroll or compute in isolation, it's the total cost of ownership for the capability: engineering time, compute, any third party data or tooling cost, rolled up against the stage split. That TCO figure is what supports the capitalized balance, what justifies whether the project was worth doing, and what a future project can be benchmarked against when someone asks how much a similar capability should cost to build.
Keep the TCO rollup at the project level, not the team level, since a single team often runs several AI projects at once and a team-level number blends a cheap, mostly finished project with an expensive, still-exploratory one in a way that's useless for judging either.
A workable tracking setup covers these points:
- Have engineers log time against both project and stage, using the project management tool you already run.
- Separate exploratory work, where you are deciding whether an approach will work, from committed build work.
- Tag training and evaluation jobs with project and stage when they are submitted, so compute cost splits the same way payroll does.
- Roll headcount, compute and any third party data or tooling into one total cost of ownership figure by stage.
- Review logged stage allocation every quarter to catch both undercounting and overcounting.
- Track fine-tuning of a third party model as its own category, separate from training a model from scratch.
Reviewing headcount allocation quarterly
Time tracking drifts if nobody checks it. A quarterly review comparing logged stage allocation against what engineering leadership's own sense of the project's progress would suggest catches both undercounting, engineers who forget to log time, and overcounting, time logged against a project that's mostly idle while someone's actually working on something else. Neither error is malicious, but both compound if left unchecked for a full year.
The review doesn't need to be adversarial. Frame it to the team as a check on the process, not an audit of any individual's time, and it's far more likely to surface honest gaps early rather than encouraging people to quietly patch their own numbers to avoid a difficult conversation.
Treating fine-tuning of a third party model as its own category
Fine-tuning a third party foundation model and training a model from scratch look similar in a project plan and don't always belong in the same capitalization bucket. Fine-tuning usually means a much smaller compute bill and a shorter application development stage, since most of the underlying capability was already built by the model provider, not by your team. Training from scratch or substantially retraining a base model's core weights is a larger, longer effort that more plainly resembles the traditional software development ASC 350-40 was written around, though your auditor decides how the guidance applies to your facts.
Track the two separately in your stage tagged time and compute data rather than rolling them into one undifferentiated AI project bucket, even when the same team works on both in the same quarter. A blended TCO figure that combines a light fine-tuning pass with a heavy from-scratch training run tells you less about either one than two separate figures would, and makes it harder to compare the cost of a future fine-tuning project against this one specifically.
This distinction also matters if you ever need to explain the capitalized balance to an auditor or investor: a much smaller capitalized figure for a fine-tuning project is expected and doesn't need defending the way an unusually small figure for a from-scratch training effort would.
What Good Looks Like
Good looks like a TCO figure by project, split by stage, that both engineering leadership and finance would sign off on as accurate.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Tracking time against project and capitalization stage is easiest when it happens inside the project management tool engineers already use daily, and a tool like ClickUp can carry that stage tag without adding a separate system.
Writing the stage-tagging habit into a documented, repeatable checklist keeps it consistent as new engineers join a project, which is exactly the kind of branching SOP work Process Street is built for.
Frequently Asked Questions
Does time spent on a failed training run still count as capitalizable?
Yes, if the run was part of committed application development, the time and compute cost of a failed attempt generally still counts, much like a bug fix during software development. If the run was still testing whether the approach works at all, it is preliminary stage work and gets expensed.
How granular should stage tagging be?
Granular enough to reconstruct the split between preliminary and application development stages by project, without becoming so detailed that engineers stop doing it consistently. Project and stage, tracked at the ticket or task level in whatever tool you already use, is usually the right level of detail.
Should contractor and offshore team time be tracked the same way?
Yes, on the same basis as internal engineering time. The capitalization test is about the nature of the work, preliminary versus application development, not who's performing it, so contractor invoices and time logs need the same stage tagging as your own team's hours.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.
Build Your Own AI Inference Cost Model in Three Tabs
How to structure a spreadsheet that turns token usage into a real cost per customer, so you can see GPU and API spend before the invoice arrives.
Why AI-Native Software Runs Lower Gross Margins Than SaaS
How variable inference cost changes gross margin for an AI-native product versus classic SaaS, and how to explain the gap to a board without a red flag.
Edge vs Cloud AI Inference: When On-Device Actually Pays Off
How to find your own crossover point between on-device AI inference and a cloud API, once you count hardware, model limits, and update infrastructure.
Budgeting for an AI Security Audit Before It Surprises You
Plan the AI-specific scope of a security audit: what changes, where audit hours go, and how to budget for evidence and vendor paperwork in advance.
What an AI Infrastructure SPV Actually Commits You To
How a special purpose vehicle isolates an AI compute commitment from your balance sheet, and the governance, exit, and diligence terms worth checking first.