GenAI costs are spiralling.  Here’s how to address the challenge.

GenAI bills are climbing faster than any budget model predicted. Why agentic AI multiplies token cost, and how orchestration makes cost visible per workflow.

In my last article, I argued that framing return on investment purely around cost reduction is the wrong approach when it comes to GenAI and that the real prize is freeing your best people to do more for more clients. I stand by that.

However, there is a different cost conversation that can’t be avoided, specifically the cost of GenAI solutions themselves. The bills are already arriving, and most organisations have no idea how to manage them.

We have been here before

In the early days of Infrastructure-as-a-Service (IaaS), the economics looked straightforward: pay only for what you use, no upfront capital, no expensive hardware to maintain. Development teams adopted enthusiastically, spun up servers without asking permission, and for a while the bills stayed manageable.

Then usage scaled, shadow IT proliferated, and organisations started receiving invoices they couldn't explain, for infrastructure decisions already embedded in production. The FinOps discipline that eventually brought cloud costs under control took years to build and required a new function most organisations didn't know they needed until the problem was already expensive. AI is following the same path, faster and with higher stakes.

The cost creep nobody planned for

The first wave of GenAI adoption felt manageable because it looked like familiar software. User-licensed tools like Copilot and Gemini came with seat counts and predictable monthly lines on the IT budget. The shift happened when teams started building their own agents, making direct API calls, and embedding GenAI into existing products and platforms. At that point the per-seat cost model stopped applying.

AI providers have accelerated that shift. Anthropic, which originally sold enterprise access as an all-inclusive seat with bundled tokens, moved in November 2025 to a model where the per seat monthly fee covers access but all token consumption is billed separately on top of that. The subsidised token pool that made the original pricing predictable is gone.

OpenAI made a parallel move in April 2026. For most enterprises, total cost of ownership went up, not down, discovered at contract renewal rather than at signing. Analysis of 2.4 billion enterprise API calls by Optimum Partners found the blended cost per million tokens fell 67% between Q1 2025 and Q1 2026, however Enterprise AI bills went up anyway, because volume grew faster than any budget model anticipated.

The agentic multiplier

Under conversational AI, cost scales roughly in line with usage. Under agentic AI, a single action can trigger a chain of agent calls, each spawning further calls and consuming tokens at every step. What looks like one task on the surface can generate ten, fifty or a hundred times the token consumption of a standard prompt and response, and nobody in the business is watching that multiply in real time.

Industry analysis consistently finds agentic systems consuming between five and thirty times more tokens per task than standard conversational tools. Most leadership teams have little, if any, appreciation of this because it’s an architectural decision made in tech teams, not a visible line on a budget. By the time the bill arrives at a scale that demands explanation, the decisions driving it are already in production. AI bolted onto workflows built in a different era is already failing to deliver value. Now it’s generating cost that nobody can see either.

Who owns this, and what needs to change

Responsibility for AI cost typically sits between the CTO and the CFO in most professional services firms today, with neither fully owning it. This is not the first time organisations have needed a cross-functional discipline that didn't fit neatly into an existing org chart.

Cloud governance went through exactly the same phase. The FinOps Foundation's 2026 State of FinOps report, drawn from 1,192 practitioners managing over $83 billion of technology spend, confirms 78% of FinOps teams now report to the CTO or CIO and only 8% to the CFO. AI cost governance is being treated as a technology and operations capability rather than a finance function, which is the right instinct.

My advice for TCS, Fund Administration, accountancy and BPO firms is to focus on the outcome, not reporting lines.  Ensure you have someone with real decision authority who can see the full cost stack, connect it to client and workflow level outcomes, and influence the architectural decisions that determine the bill before it’s written. The prize for getting this right, and the cost of getting it wrong, are both significantly larger than they were with cloud.

The hybrid horizon

The more considered organisations are already moving beyond the pure consumption model. Rather than relying exclusively on API-based AI and cloud-hosted inference, they are investing in their own GPU infrastructure and on-premise or co-located compute, the AI equivalent of the hybrid cloud strategies that emerged once enterprises understood their IaaS exposure well enough to make build-versus-buy decisions.

While owning the infrastructure may bring higher upfront capital cost, it delivers predictable unit economics, removes exposure to provider repricing, and keeps sensitive client data inside a controlled environment. For regulated service providers in the TCS, Fund Administration, and accountancy space, that last point is essential.

The visibility gap that still needs to be closed

Most TCS, Fund Administration, accountancy and BPO firms have no mechanism for combining people cost and AI cost into a single view of what it’s actually costing to serve a client, run a case, or resolve an exception. A monthly API invoice doesn't know which client's work triggered which model calls. A time recording system doesn't know which steps were automated and which required human intervention. Nobody is joining those datasets together because nobody has built the layer that would make it possible.

This is where orchestration is critical. It’s the missing layer that decides which tool handles which step and who owns the handoff, and is also the only point in a reimagined workflow where cost can be tracked at the granularity that actually matters. Without it, cost governance is a retrospective exercise on data never designed to answer the question. With it, cost visibility at the per-client, per-case level becomes a natural output of how work is run. The right unit of measurement is not cost per token. It’s cost per workflow, the only number that can sit alongside what the same work cost before AI and produce something a board can act on.

So where do we go next?

As AI embeds deeper into client workflows, two questions are going to force themselves onto the agenda of every TCS, Fund Administration, accountancy and BPO firm.

  1. Do your clients allow you to use their data to train, refine and improve the AI running inside their engagements?
  2. Do you understand the intellectual property (IP) status of the outputs your AI is producing?

These are not technology questions. They are legal, commercial and reputational.  Most firms in this sector haven’t started asking them yet. That’s the subject of the next article.