Ch. du Vernay 14a
CH-Vaud
1196 Gland
info@neumarz.com
+41.21.561.34.96
Back

Token Prices Fell 67%. Three in Four AI Budgets Broke Anyway.

TLDR: The price of a token is falling and the token bill is rising, which means the budgeting variable that matters inside a bank or a fund is no longer what a token costs but how many tokens a task consumes.

Token expense has arrived on the earnings calls

For two years, AI cost inside financial institutions was a technology line item that nobody outside the technology function interrogated. That changed this quarter. Speaking after Goldman Sachs reported, David Solomon acknowledged that “there’s a lot of talk about token spend and the cost of the technology”, and warned that the build-out “won’t be without bumps and re-calibrations as people can understand just what the ultimate demand is for this technology in enterprises.”

Morgan Stanley chief financial officer Sharon Yeshaya was more precise about the trajectory on the firm’s second-quarter call. Token expense, she said, is “one topic which is not financially meaningful this year, but which may become so in the future”, and it is already absorbing management attention.

Read that carefully, because it is the whole problem in a sentence. A cost that is immaterial today and material tomorrow is precisely the kind of cost that escapes budget discipline until it is too large to unwind quietly.

The price collapse is real, and it is a trap

Every headline number supports optimism. The blended cost of a million tokens fell 67% year over year. Frontier output pricing has dropped roughly 94.5% since March 2023 on public benchmarking of inference prices. The spread between tiers is now enormous: a commodity model runs at cents per million input tokens while a frontier reasoning model runs at several dollars, with output priced multiples higher again.

And yet, in the same survey period, 73% of organisations reported that AI costs exceeded their original budget planning, while the share of teams actively managing AI spend climbed to 98% from 31% two years earlier. Falling unit prices, rising bills, and near-universal cost management that is still not containing the outcome.

The mechanism is elastic demand. When the marginal cost of an inference falls, organisations do not bank the saving. They lengthen reasoning chains, add tool calls, widen context windows and hand work to agents that call other agents. AT&T has publicly reported scaling from roughly 8 billion to 27 billion tokens per day after deploying multi-agent systems. Google has reported processing on the order of 1.3 quadrillion tokens per month, up roughly 130-fold year over year.

The variable that moves the bill is tokens per task

This is where a quantitatively trained reader should focus. Spend is price multiplied by volume. Price is falling on a reasonably predictable curve and is visible on a rate card. Volume is not, because volume is not a function of user headcount. It is a function of architecture.

Reasoning and agentic workloads consume five to thirty times more tokens per task than the equivalent chat interaction. A single query routed through a retrieval pipeline with a reasoning model and three tool calls can consume one to two orders of magnitude more tokens than the same question asked directly of a smaller model. Reasoning models generate long internal chains of thought that are billed as output tokens, so the consumption profile diverges sharply from what the interface appears to show.

The result is that token consumption is non-linear with respect to user-facing activity. That single property explains most of the forecast error in the market. A bank that models AI cost as seats multiplied by a price per seat has built a linear model of a non-linear system, and the error compounds in the direction of overspend every time an engineering team upgrades a model or gives an agent more autonomy.

Tokens are also not homogeneous, which defeats naive volume tracking. Serving economics sit on a frontier between throughput and interactivity, and the useful measure is goodput: output that meets a defined latency and speed objective rather than raw tokens produced.

Token class Profile Where it belongs
Bulk High throughput, low per-user speed, cheapest to serve Overnight document processing, embeddings, batch summarisation
Goldilocks Moderate interactivity at near-optimal throughput Most internal analyst and client-service applications
Premium low-latency High per-user speed, lower aggregate throughput, higher unit cost Voice agents and workflows where response time gates productivity
Reasoning Looks like a standard call; generates many internal tokens per returned token Genuinely hard analytical problems only; the main source of spend surprise
Framework adapted from FinOps Foundation token economics and published inference benchmarking (2026).

The subsidy phase has ended

The assumption quietly embedded in most multi-year AI budgets is that next year’s price cuts will absorb this year’s volume growth. That assumption is now unsafe, because providers have started repricing toward cost recovery.

Anthropic’s April 2026 enterprise transition is the clearest example. The company moved enterprise customers to a seat fee plus pre-committed token consumption, with no included usage cushion, and separated flat-rate subscriptions from third-party agent harnesses that analysts estimated were consuming several times more compute per dollar than the subscription was priced to support. OpenAI’s ChatGPT product lead made the same point more bluntly, observing that an unlimited AI plan is structurally similar to an unlimited electricity plan and does not work.

The practical consequence for a treasury or CFO function is that buyers must now forecast compute demand the way they learned, painfully, to forecast cloud demand a decade ago. Per-token list prices continue to drift down, but the declines are concentrated in commodity tiers, while the tiers that agentic workloads actually require are holding firmer.

What disciplined institutions are doing

Yeshaya described Morgan Stanley’s approach in terms any operator would recognise as ordinary cost engineering: “use the right model for the right purpose, be smart about open source where appropriate”. Her worked example is instructive precisely because it is mundane. Summarising an analyst report is a real workload with real value, and “you really don’t need the latest cutting-edge incredibly expensive model” to do it.

That is model routing, and it is the single highest-return control available. Around it sit the rest: caching repeated context, batching anything that does not need to be interactive, capping agent recursion depth, and metering goodput rather than raw token counts so that a fast token and a slow token are not treated as the same economic good.

The accounting matters as much as the engineering. BCG’s framing in its work on managing the token meter is that token spend does not belong in a single bucket. Tokens that build reusable capability behave like investment. Tokens that run internal work are operating expense. Tokens consumed inside a product or a client interaction are cost of goods sold. An institution that books all three as IT overhead cannot tell which of its AI programmes is creating value and which is quietly eroding a margin.

The final control is attribution. Token consumption charged back to the desk that generated it changes behaviour in a way that a central budget never does, for the same reason that metered utilities are used more carefully than flat-rate ones.

Why this is a margin question, not a technology question

Here is the part that gets least attention and matters most to owners and allocators.

Jamie Dimon rejected the idea that AI efficiency simply accrues to the firm that deploys it. “We can’t just say, ‘Oh, it’s going to increase our margins, and we’re going to keep that’. If that were true, our margins would be 80% today because of computerization over the last 20 years.” He went further: “you don’t uniquely benefit from AI. The ultimate beneficiary of AI will be our customers.”

Competition passes the productivity gain through to the client. What stays on the income statement is the cost. For an asset manager or a fund, that cost is migrating out of one-off project capital and into recurring cost of service, landing on a fee base that is already compressing structurally. A technology decision therefore becomes a question about the durability of operating margin, which is the input that drives the multiple a buyer will pay.

Nor is the input environment about to soften. Morgan Stanley chief executive Ted Pick noted that data-centre capital expenditure forecast at $575 billion for 2026 is arriving closer to $850 billion, with 2027 revised from roughly $700 billion to $1.3 trillion, and estimated that the industry is only 10% to 15% through the investment cycle. Suppliers building at that pace will eventually need to earn a return on it.

What to do about it

If you run an investment manager or a fund. Instrument cost per task before you scale a use case, not after. Set an explicit token budget per workflow with a threshold at which the workflow is re-engineered or switched off, and route each task to the cheapest tier that clears the quality bar. Classify token spend as investment, operating expense or cost of goods sold, and charge it back to the desk that consumes it. Dimon’s own framing is a useful discipline: JPMorgan runs close to a thousand AI use cases, of which he counts roughly 50 as genuinely important.

If you are diligencing or allocating to one. Stop counting use cases. Ask for cost per task, the tier mix behind it, and the trend in tokens per task over the last two quarters. A firm with a thousand pilots and no unit economics has built a cost centre with good public relations. A firm that can quote its cost per completed workflow, and show it falling while volume rises, has built something that survives the repricing.

The institutions that come through this cycle intact will not be the ones that spent least on inference. They will be the ones that always knew what a unit of it cost them, and what it bought.

References

  1. diginomica. “Tokenomics — the direction of investment travel from Goldman Sachs, JP Morgan Chase, and Morgan Stanley” (Stuart Lauchlan, 17 July 2026). https://diginomica.com/tokenomics-direction-investment-travel-goldman-sachs-jp-morgan-chase-and-morgan-stanley
  2. FinOps Foundation. “Token Economics: The Atomic Unit of AI Value” (J.R. Storment). https://www.finops.org/insights/token-economics-the-atomic-unit-of-ai-value/
  3. FinOps Foundation. State of FinOps 2026 Report. https://data.finops.org/
  4. Epoch AI. “LLM inference prices have fallen rapidly but unequally across tasks.” https://epoch.ai/data-insights/llm-inference-price-trends
  5. Boston Consulting Group. “Return on AI: How CFOs and CIOs Can Manage the Token Meter” (2026). https://www.bcg.com/publications/2026/managing-ai-token-costs
  6. Morgan Stanley. Q2 2026 earnings call transcript (15 July 2026). https://www.fool.com/earnings/call-transcripts/2026/07/15/morgan-stanley-ms-q2-2026-earnings-call-transcript/
  7. Deloitte Insights. “AI tokens: How to navigate AI’s new spend dynamics.” https://www.deloitte.com/us/en/insights/topics/emerging-technologies/ai-tokens-how-to-navigate-spend-dynamics.html

Leave a Reply

Your email address will not be published. Required fields are marked *

Legal Disclaimer

Investments involve a high degree of risk.

Investors should carefully consider the risks described & all other information in the investor agreement provided to you before deciding whether to invest in the proposal.

Participation and investment are speculative activities that involve a high degree of financial risk. Certain risk factors that should be considered in assessing an investment in Neumarz (an Allegory Capital & Kainjoo SA brand) and its activities include, but are limited to, those set out below.

Any one or more of these risks could have a material adverse effect on the value of any investment in Neumarz, an Allegory Capital & Kainjoo SA brand, and the business, financial position or operating results of Neumarz, an Allegory Capital & Kainjoo SA brand.

An investor may lose all or part of his or her investment in Neumarz, an Allegory Capital & Kainjoo SA brand.

Additional risks and uncertainties not currently known to the officers and directors of Neumarz, an Allegory Capital & Kainjoo SA brand, may also adversely affect current activities. The information below is not an exhaustive summary of the risks affecting Neumarz, an Allegory Capital & Kainjoo SA brand. It is not intended to be presented in any assumed order of priority. The hazards relating to the business of Neumarz, an Allegory Capital & Kainjoo SA brand, include, among other things:

(a) there are no assurances that Allegory Capital  or Kainjoo SA will earn profits in the future or that profitability will be sustained;

(b) there are no assurances that Allegory Capital  or Kainjoo SA will have access to sufficient funding for future operations or to fulfil its obligations under current agreements;

(c) risks that are inherent to investments, including (i) the valuation of these assets is subject to significant volatility, (ii) the regulatory regime governing venture capital investments is uncertain, and new regulations or policies may materially adversely affect the development of investment, (iii) the further development of businesses are subject to a variety of factors that are difficult to evaluate, (iv) companies and intangible assets are at risk of security breaches, (v) potential loss or destruction due to cybersecurity threats, and (vi) risks of an illiquid market for assets;

(d) difficulty in valuing investments;

(e) Allegory Capital and Kainjoo SA have limited operating history, and there is no assurance that the investments of Allegory Capital will be profitable;

(f) Allegory Capital has not generated revenues to date, and there can be no guaranteed return on its investments;

(g) directors, officers and key employees may resign from Allegory Capital ;

(h) Allegory Capital is unable to insure against every risk to which it is exposed;

(i) Laws relating to the business of Allegory Capital may be changed in a manner which adversely affects Allegory Capital ;

(j) Allegory Capital may invest in entities with no operating history, making evaluating such entities complex and

(k) Risk related to foreign exchange rates.

Additional risks and uncertainties not presently known to Allegory Capital and Kainjoo SA or that it currently deems immaterial may substantially affect its business, financial condition, valuations, trading performance and prospects.

Potential investors are accordingly advised to consult an independent financial adviser who specializes in advising on investments of this kind before making any investment decisions in Allegory Capital or Kainjoo SA.

A prospective investor should consider whether an investment in Allegory Capital or Kainjoo SA is suitable in light of his or her circumstances and the available financial resources.