AI Token Spend Is Becoming a GAAP Problem: COGS, OpEx, or Capitalizable?

TokenCheat Team

7/2/2026

#finops#accounting#cost-analysis#governance
AI Token Spend Is Becoming a GAAP Problem: COGS, OpEx, or Capitalizable?

This is analysis, not accounting or tax advice. Talk to your auditor and tax advisor before changing how you classify anything.

Every month, a bill arrives that looks like this:

Anthropic PBC ............... $47,231.18

One line. Behind it: tokens burned building your next product, tokens burned maintaining legacy code, tokens burned serving customers in production, tokens burned on a research spike that went nowhere. Under GAAP, those four uses have different accounting treatments — and almost no engineering organization can tell them apart, because the attribution data doesn't exist.

What the tokens didLikely treatmentWhere it lands
Built internal-use software (development stage)Capitalizable candidate (ASC 350-40)Balance sheet, then amortization
Built software sold to customersASC 985-20 analysis (technological feasibility)Depends on stage
Served customers in production (inference)Cost of revenueGross margin
Maintained existing codeOperating expense (typically R&D)OpEx
Research spike that went nowhereOperating expenseOpEx

(Directional summary, not advice — your auditor decides where the lines fall.)

That was a rounding error when AI spend was a few Copilot seats. At $500–2,000 per developer per month — the range usage-based tools like Claude Code commonly reach for heavy users — a 50-developer organization is looking at $300K–1.2M a year (straightforward arithmetic, your numbers will vary). That's past the point where "dump it in software subscriptions" survives an audit committee's attention.

The three classification questions

1. Is it capitalizable? (ASC 350-40)

For internal-use software, GAAP has long required capitalizing costs incurred during the application development stage — explicitly including external direct costs of materials and services consumed in developing the software, alongside payroll for developers working directly on it. Tokens consumed by a coding agent building your internal platform are a very natural reading of "external direct costs of services consumed in development."

Two implications most teams haven't confronted:

  • If your developers' AI spend is materially producing capitalizable software, expensing all of it may understate assets and EBITDA relative to the treatment you're already applying to those same developers' salaries.
  • FASB has modernized this area: ASU 2025-06 (Targeted Improvements to the Accounting for Internal-Use Software) retires the rigid project-stage model in favor of a probable-to-complete threshold — a change built for how agile and AI-assisted development actually works. It takes effect for fiscal years beginning after December 15, 2027, with early adoption permitted. The practical effect: capitalization decisions will hinge on project-level judgment, which requires project-level cost data.

(Software sold to customers follows ASC 985-20 instead, with its technological-feasibility threshold — a different analysis, same attribution prerequisite.)

2. Is it COGS or OpEx?

If model inference runs inside your product serving customers, that spend belongs in cost of revenue — it hits gross margin, the number investors scrutinize hardest for AI-native companies. If it's development tooling, it's operating expense (typically R&D). Most AI-native companies consume both through the same provider accounts. Splitting a single invoice between COGS and R&D without usage-level attribution is guesswork, and the error lands in the most-watched line of the income statement.

3. What about tax?

US tax treatment of research and experimental expenditures has whipsawed: TCJA-era rules required capitalizing and amortizing R&E (software development is explicitly R&E) starting in 2022; 2025 legislation restored immediate deduction for domestic R&E for tax years beginning after 2024, while foreign R&E generally still amortizes over 15 years. Add R&D-credit substantiation, and the common thread is the same: the ability to tie AI spend to specific projects and activities, with evidence.

The real problem is attribution, not policy

Note what all three questions have in common. None of them is hard because the accounting guidance is unclear. They're hard because the data doesn't exist:

  • Provider invoices arrive as one line per account — not per project, per repo, per cost center, or per development phase.
  • Seat-based tools (Cursor, Copilot) bury token economics inside a flat fee with opaque overage behavior — you can't see consumption at all.
  • Nothing in the default tooling records what the tokens were spent doing — new feature development (capitalizable candidate), maintenance (expense), production inference (COGS), or exploration (expense).

Compare how the industry treats cloud spend: a decade of FinOps practice produced tagging, showback, chargeback, and per-service cost allocation precisely because finance demanded it. AI token spend is at the pre-tagging stage of that same curve — material enough to matter, instrumented like it doesn't.

What defensible attribution looks like

If your auditor, controller, or board ever asks, the evidence trail they'll want is roughly:

  1. Spend by repo/project — which codebase consumed the tokens
  2. Spend by activity — development vs. maintenance vs. production inference (separate API keys per use is the crude-but-effective start)
  3. Spend by team/cost center — for chargeback and departmental budgets
  4. Per-unit economics — cost per PR, per feature, per task: the bridge between token telemetry and project accounting
  5. A consistent method, documented — auditors forgive estimation; they don't forgive improvisation

Practical first steps that cost almost nothing today: separate production API keys from development keys (COGS/OpEx split solved at the account level); keep a mapping of repos to projects and cost centers; capture usage telemetry per developer and per repo; and raise the classification question with your auditor before the spend is material, not after.

An honest note on materiality

If your organization spends four figures a month on AI coding, none of this justifies process. Auditors care about material misstatements, and today most AI coding budgets are still small next to the payroll of the developers using them. But the trajectory is one-directional — spend per developer is rising, agentic workflows multiply token consumption, and the companies that instrumented cloud costs early had a much easier decade than the ones that retrofitted. The cheap move is capturing the attribution data now, while the policy question is still optional.

Where TokenCheat fits

TokenCheat's telemetry layer records exactly the substrate this problem needs: usage snapshots by organization, member, repo, and tool, with cost roll-ups like cost per PR. The free stack audit shows you the per-session anatomy of your spend in under a minute — and if your finance team is already asking these questions, that's precisely the conversation our team audit exists for.

Again: not accounting or tax advice. Bring your auditor a coffee and this article instead.