A CFO I worked with last year had exactly one question about a proposed AI program, and it was not about accuracy or headcount. She asked how long the thing would last. Her team had already capitalized a chatbot build in 2023, put it on a five-year amortization schedule, and then watched the underlying model get deprecated fourteen months later. The software was still on the books. The capability was not.
That is the actual finance problem with AI right now. Not whether it works. Whether the accounting you use to fund it describes anything real.
The $690B question your board is already asking
Aggregate AI capital expenditure is projected to reach roughly $690 billion in 2026, according to Futurum's infrastructure analysis, with hyperscalers carrying most of it. Allianz Research has called the cycle war-proof for now while flagging that the revenue base underneath it is thinner than the spend implies. The BIS reached a similar place in its 2026 Annual Economic Report, treating AI investment as a genuine productivity bet with an unresolved timing problem.
None of that is your problem directly. You are not building data centers. But it changes the conversation in your boardroom in one specific way: directors have read the bubble coverage, and they now ask harder questions about depreciation than they did about your ERP migration.
The question that keeps landing is some version of "if the models get better and cheaper every six months, why are we writing a three-year check?" It deserves a real answer, and the real answer is that most of what you are funding is not the model.
What actually depreciates
The mistake almost everyone makes is treating "the AI system" as one asset with one useful life. It is at least three assets with wildly different decay rates, and once you separate them the budgeting gets straightforward.
| Layer | What it includes | Useful life | Survives a model swap? |
|---|---|---|---|
| Durable | Data pipelines, schema contracts, permission model, audit logging, eval harness, system integrations | 3-5 years | Yes |
| Semi-durable | Orchestration logic, tool definitions, guardrails, retrieval index design, human handoff paths | 18-36 months | Mostly |
| Perishable | Model selection, fine-tunes, prompt scaffolds tuned to a version, provider-specific optimizations | 6-18 months | No |
| Consumption | Inference tokens, hosting, human review time | Ongoing | N/A |
The eval harness is the clearest example of a durable asset that finance teams routinely misclassify as a project cost. A good eval suite is the thing that lets you swap providers in a week instead of a quarter. It is the highest-return line item in the entire build and it has nothing to do with any specific model.
The perishable layer is where the honesty problem lives. If you spent $180,000 building a fine-tune and a set of prompts calibrated to a model that gets retired next spring, that money bought you fourteen months of capability. Amortizing it over five years produces a balance sheet that shows an asset you cannot use.
Public filings show the same tension at much larger scale. Hyperscalers have moved server useful lives in both directions over the past three years, extending them to six years when utilization looked durable and then trimming them back as AI-specific hardware turned over faster. Amazon shortened the assumed life on a subset of servers and networking equipment effective January 2025 and disclosed a material hit to operating income as a result. You can pull the exact language from SEC full-text search across 10-K filings. The point for your own model is that useful-life assumptions are a judgment, they are disclosed, and they get revisited when reality disagrees.
What the rules actually let you capitalize
US GAAP treats internally developed software under ASC 350-40, and the shape of it is familiar to anyone who has funded a platform build. Preliminary planning is expensed. Application development is capitalized. Post-implementation operation and maintenance is expensed. FASB issued targeted improvements in late 2025 that move away from the rigid project-stage framework toward a probability-of-completion threshold, which helps AI work specifically, because AI projects rarely proceed in clean sequential stages.
A practical mapping for an agent build:
- Capitalize: integration code, data pipeline construction, orchestration and tool layer, eval harness construction, permission and audit infrastructure, configuration and testing during the build.
- Expense: discovery and use case selection, data cleansing (explicitly excluded), training and change management, prompt iteration after go-live, ongoing evaluation runs.
- Never capitalize: inference spend. Tokens are consumption. Some teams try to argue that inference used during development is a build cost, and in narrow cases that holds, but a recurring monthly bill is opex no matter how the invoice is titled.
If you are consuming a hosted model rather than running your own, the arrangement is a service contract, and implementation costs follow the cloud computing guidance in ASU 2018-15: capitalize qualifying implementation costs and amortize them over the term of the hosting arrangement, including reasonably certain renewals. That term is your ceiling. A twelve-month API agreement does not support a sixty-month amortization schedule.
On the tax side, domestic research expensing was restored for tax years beginning after December 31, 2024, which reversed the five-year Section 174 capitalization that made AI R&D painfully expensive in 2022 through 2024. Your book treatment and your tax treatment will diverge. Say so in the memo before someone finds it.
A multi-year AI budget model that survives a deprecation
Here is the structure I use. Three years, four lines, one refresh assumption baked in from the start.
| Line | Year 1 | Year 2 | Year 3 |
|---|---|---|---|
| Build (capitalized) | $540K | $120K | $180K |
| Build (expensed) | $360K | $80K | $90K |
| Run (inference, hosting, review) | $180K | $310K | $420K |
| Model refresh reserve | $0 | $150K | $150K |
| Total cash | $1.08M | $660K | $840K |
Three things about this table that finance people notice immediately.
The run line grows. Per-token prices for a given capability level have fallen steeply, and Epoch AI's cost tracking documents the trend well. But unit cost falling does not mean spend falling. Volume rises because the system works, and per-task token consumption rises because reasoning models think longer. Budget volume and unit cost as separate lines so you can hold each accountable. If unit cost is flat while volume triples, that is a win. If both rise, you have a routing problem.
There is a standing refresh reserve. Not a contingency. A scheduled line, sized at roughly 20 percent of the original build, that funds re-tuning and re-validating against a new model generation. Teams that skip this line end up either frozen on a deprecated model or raiding the run budget mid-year.
The capitalized share drops fast. Year 1 is 60 percent capitalizable because you are building durable infrastructure. By year 3 you are mostly operating. A program where the capitalized share stays high in year 3 is either genuinely expanding scope or quietly reclassifying maintenance, and your auditors will ask which.
For the return side of the model, you need a baseline that predates the system. Measuring cycle time and volume before you automate is unglamorous work that determines whether your ROI number means anything. McKinsey's State of AI survey has consistently found that the share of firms reporting EBIT impact from AI is far smaller than the share deploying it, and a large part of that gap is measurement, not capability. The Anthropic Economic Index is useful as an outside-in check on which task categories are actually absorbing AI usage, which helps sanity-test whether your projected volume is plausible.
Fund tranches against stability gates, not maturity stages
The standard consulting framing puts AI programs on a maturity curve: pilot, scale, transform. It is a bad basis for releasing money because maturity stages are self-assessed. A team can declare itself "scaling" while error rates are flat and nobody trusts the output.
Gate each tranche on measured thresholds that have to hold for a window:
- Gate 1 (release build tranche 2): end-to-end task completion above target on a held-out eval set, with the eval set built from real failure cases rather than synthetic examples.
- Gate 2 (release scale tranche): escalation rate flat or falling over 30 consecutive days at production volume, and cost per completed task inside the modeled band.
- Gate 3 (release expansion tranche): the workflow has run through at least one model version change without a rebuild.
That third gate is the one nobody sets and everybody needs. It is the empirical proof that you built durable assets rather than a scaffold around a specific model. A system that survives a provider swap has earned a multi-year amortization schedule. One that has never been tested has not. The same logic governs when to scope the next workflow: stability in the current one is the precondition, not enthusiasm about the next one.
Academic work is starting to formalize this distinction. A multi-method evaluation of the AI investment cycle separates infrastructure buildout from application-layer value capture and finds the two decoupling on very different timelines. Your internal model should make the same separation.
The three numbers your board memo needs
Strip out the architecture diagrams. A board wants three figures and a range.
1. Fully loaded cost per unit of work. Not cost per token. Cost per resolved ticket, per reviewed contract, per qualified lead, including human review time and the amortized build. Compare it to the current fully loaded cost of the same unit. If you cannot produce this number, you do not have a business case, you have a budget request. Measuring the outcome rather than the prompt is the whole discipline here.
2. Durable share of spend. What percentage of the multi-year total buys assets that survive a model change. Below 30 percent means you are renting a capability with extra steps. Above 70 percent means someone is capitalizing things they should not be.
3. Model swap cost. The engineering cost and calendar time to move the workflow to a different provider. This is the single most informative number in the entire package and almost nobody produces it. A program with a swap cost of two engineer-weeks has real optionality and deserves a longer funding horizon. A program with a swap cost of four months is a bet on one vendor's roadmap, and it should be funded and disclosed as one. a16z's work on compute economics makes the structural version of this argument, and it applies at every scale.
Present a range on all three, with the pessimistic case assuming a model deprecation inside eighteen months. Boards trust ranges. They stop trusting point estimates the first time one misses.
How OpenNash Can Help
Most of the AI programs that fail their second budget review did not fail technically. They failed because nobody could show which part of the spend was an asset and which part was rent.
OpenNash builds production AI systems with that separation designed in from the audit stage: durable pipelines, evals, and integrations built to outlive any specific model, with the model-dependent layer kept thin and cheap to replace. Deployment includes full ownership handoff, documentation, and the eval harness, which is what makes the swap-cost number small enough to defend.
Platform tools are the right answer when your workflow is standard and your volume is modest, and waiting one more quarter is the right answer when you have no baseline measurement. Custom build makes sense when the workflow is specific to how your business actually runs, when auditability matters to a regulator or a customer, and when you need to own the asset rather than lease access to it. Book a call to map this budget structure to your workflow.
The next model generation will land before your amortization schedule ends. Write the schedule as if you already know that.