Connect with us

NEWS

SAP Puts Token Spend on the Same Line as Hiring

SAP now charges AI token spend to department books, trading a 30 percent R&D lift for slower hiring while most firms still cannot see the bill.

Published

on

SAP now books AI token spend against the same department budgets that used to hold hiring, after an average 30 percent lift in R&D productivity. In a September 2026 note, chief controlling officer Lukas Deutsch and David Imbert, chief marketing officer for SAP Financial Management, said the meter has to produce a return, not just a smaller invoice.

Chief executive Christian Klein had already described the trade on the July 23 results call, when SAP posted a current cloud backlog of €22.9 billion. Tokens, he said, now sit beside headcount in one manager budget, and the company will miss the hiring plan it set at the start of the year.

SAP Already Trades Tokens for Headcount

Klein told investors the heaviest token use is in R&D, where productivity is up by an average of 30 percent, so extra people are no longer the default. Cost has been rising faster than headcount, which he called the token effect, because SAP charges the tokens to the functions that use them. A few expensive specialists still join. The broader recruiting machine does not.

You have a head count budget, you have third party budget, and you have tokens. So our managers in R&D, in the go-to-market space, they have one budget and they have to manage that also according to the productivity assumptions we reflected in the budget.

Christian Klein, CEO, SAP Q2 2026 earnings call

Cloud revenue in the quarter was €6.3 billion, up 24 percent at constant currencies, and total revenue was €9.9 billion, up 11 percent. Non-IFRS operating profit rose 9 percent at constant currencies. Klein still needs an 80 to 90 percent revenue-to-cost ratio, and he is trying to hit it by balancing tokens against people rather than by adding both.

About 4,000 people inside SAP already use Joule Work, the company’s agent interface, which it said would reach customers in the third quarter of 2026. That is the internal proof behind the finance argument: the bill is real, the output is showing up in R&D, and hiring is the line that gives way.

Why Token Costs Now Sit on Department Books

Deutsch and Imbert wrote that a single company-wide AI budget makes early trials easy and then breaks. The people burning tokens never see the cost, so they never have to defend it. Central pots also hide which teams, which jobs, and which tools are driving the meter.

They are not asking finance to invoice every model call to the last decimal. They want tokens treated like software, outside services, and labor: someone owns the resource, and that person can say what it is supposed to return. Once SAP pushed token costs into business areas, they said the questions changed from how much was spent to what the spend produced, and whether more money belonged there.

Klein made the same design choice in public. Managers in R&D and go-to-market do not get a free AI overlay on top of payroll. They get one envelope. If tokens rise, something else in that envelope has to move, which is why the hiring plan is already being cut. Philipp Herzig, SAP’s chief AI officer, told a Bank of America conference in June that customers “really don’t like tokens; they like business outcomes,” which is why SAP prices much of its own AI on value rather than passing the meter through.

That split matters inside SAP too. The company can switch models on its AI platform for the best price-to-outcome ratio, Klein said, and it is not locked to one frontier vendor. Routing is how a department stays inside a budget without banning the work.

93 Percent Blew Past the AI Budget

Most firms are not yet having that conversation because they cannot see the bill in pieces. A McKinsey survey published on July 20 found that 93 percent exceeded their AI budgets, even though 62 percent have moved past experiments into live use. Spend jumped nearly fourfold as programs went from isolated pilots to company-wide use, and a majority expected outlays to rise at least 25 percent over the next 12 months.

McKinsey’s May 2026 Enterprise AI FinOps survey also found that only 20 to 25 percent of companies have mature AI FinOps. In McKinsey’s field work, 20 to 30 percent of AI spend is often unaccounted for because it is scattered across copilots, foundation-model contracts, software features, APIs, and lab accounts. The same task can burn up to 30 times as many tokens depending on the model, the agent chain, and how sloppy the prompt is.

A Harness survey of 700 engineering and FinOps leaders in May and June 2026, released as the 2026 State of AI in FinOps, found that 52 percent named no cost owner. Responsibility sat in pieces across engineering, FinOps, finance, and IT. Patrick Brogan, director of FinOps advisory at Harness, said AI spend had moved from a line that occasionally surprises people to a category that regularly does.

THE AI SPEND BLIND SPOTS

Blind spot Share Survey
Unexpected AI bill spike in the past year 72% Harness, 700 leaders
Could explain a doubled bill within hours 20% Harness
Mature AI FinOps practices in place 20-25% McKinsey, May 2026
FinOps practitioners now managing AI spend 98% FinOps Foundation, 2026

Harness respondents put wasted AI spend at 26 percent. For the one in five organizations already at $1 million a month, that is $260,000 with no measured return. Seventy-three percent said they had cost policies, yet only 13 percent had basic visibility, and only 26 percent had a solid way to measure business value. Fifty-six percent said forecasting was guesswork. Fifty-seven percent said their company actively pushed “tokenmaxxing,” more use whether or not it paid off.

The FinOps Foundation’s 2026 State of FinOps report, with 1,192 responses, said 98 percent of practitioners now manage AI spend, up from 63 percent in 2025. McKinsey’s later State of AI 2026 survey found that about 20 percent of organizations were already limiting AI use because of operating costs, including tokens. The control function arrived after the invoice.

Faster Pull Requests Are Worth the Tokens

Deutsch and Imbert warned against using cost per token as the scorecard. A tool that speeds software delivery, cuts repeat work, or improves service will show a large token line, and cutting that line because it is visible can destroy more value than it saves. After SAP rolled out AI developer tools, they recorded a mid-double-digit percentage rise in pull-request merge rates, which they read as work moving faster, not as a bill to crush.

SAP’S OWN VALUE PAIR

  • Merge rate: AI developer tools produced a mid-double-digit percentage rise in pull requests merged.
  • R&D load: Token use is highest in R&D, the same place Klein tied to the 30 percent productivity gain.
  • Internal bench: 4,000 employees were on Joule Work before the customer launch window in the third quarter of 2026.
  • Target ratio: Klein is balancing tokens and headcount to hold an 80 to 90 percent revenue-to-cost ratio.

Finance, they argued, needs cost metrics sitting next to value metrics. Governance that only hunts a cheaper token will starve the jobs that are working. McKinsey’s own fieldwork says companies that manage consumption with some care can take 20 to 30 percent out of AI costs, and about a third of surveyed firms had already banked savings in that range in the prior three months, often by routing simple work off frontier models and by caching repeated prompts, which can cut repeated input-token costs by up to about 90 percent.

The unit that matters is the finished job, a merged change, a closed ticket, a handled case, not the token. Teams whose bills rise while cost per result falls are doing what the budget was for. A hard cap cannot tell that team from one that is wasting money unless someone has tagged the work.

Power Users, Wrong Models, and Tool Sprawl

When SAP inspected its own use, three patterns explained the outsized bills. Those are the places Deutsch and Imbert say controls should cut, not across the whole user base.

THREE ROOT CAUSES OF TOKEN WASTE

  • Power users and agents: A few people and automated agents can dominate the meter, especially when loops retry, call tools, and keep talking to the model.
  • Model mismatch: Premium models get used for jobs a smaller model could finish, because the default is whatever the developer already has open.
  • Tool sprawl: Overlapping assistants, review bots, and lab accounts bill the same work several times and scatter the invoice.

Agentic workflows make the first pattern worse. EY’s comparison of customer-service AI put a 2023 linear query at $0.04 and a 2026 orchestrated loop of tools, reasoning, and retries at $1.20, about 30 times higher per interaction. Gartner has estimated that more than 40 percent of agentic AI projects will be canceled by the end of 2027 on cost, fuzzy value, or weak risk controls.

Handing every employee a general assistant is a reliable way to mint orphan agents and a bill nobody can map to a result. In finance especially, the cheaper working pattern is mostly ordinary software, with the model called only where judgment is required, plus a named owner when the meter runs. SAP’s answer was narrower: token caps to stop runaway jobs, model routing so the task gets the cheapest model that can do it, and fewer tools so spend concentrates where use is high enough to justify it.

Someone Has to Own the Token Line

Visibility without an owner is a report. Deutsch called ownership the consequential insight, more than any forecast tweak. Chargeback is how that owner feels the cost. Showback tells a team what it used. Chargeback puts the same number on the cost center that already pays for people, which is when the ROI question appears without a slide.

Large software shops already run a version of this on their own AI credits, with chargeback aligned with P&L ownership so a cost-center owner can lift or hold the cap. User-level budgets are the hard stop; company and cost-center budgets are the warning layer unless someone configures them to cut access. Shared keys make the bill anonymous. Per-project keys make it a person.

Less than half of Harness respondents, 45 percent, said they understood the cost of the AI features they build. If the engineer picking the model never sees a unit cost, the company is running a principal-agent problem: the people consuming tokens are not the people who answer for them. SAP’s internal fix is blunt. The R&D manager and the go-to-market manager already live inside a combined budget, so a token spike is a staffing decision in disguise.

That is the second-order effect the September note only hints at. Once tokens share an envelope with payroll, a new agent has to beat a new hire, and a noisy copilot has to beat both. Tools that cannot show a merge-rate lift, a shorter close, or a handled case lose the argument even if the model is impressive.

The Nine-Figure Risk They Chose Not to Freeze

Deutsch and Imbert said the three levers, caps, routing, and fewer tools, helped contain a triple-digit-million financial risk without slamming the brakes on adoption. The rule they kept was simple: strip waste, leave productive demand alone. They also said the work is unfinished. The next job is to put the same habits into ordinary planning and reporting, sharpen ownership, improve allocation, and forecast costs before they land.

TOKEN DECISIONS INSIDE SAP THIS YEAR

  1. July 23, 2026: Klein tells investors tokens are charged to the functions that use them and that R&D productivity is up by an average of 30 percent.
  2. July 2026: SAP reviews hiring plans for the next 12 months and for 2027, and Klein says this year’s intake will nowhere near the original plan.
  3. Third quarter of 2026: Joule Work is due to customers after 4,000 internal users, with close to 50 assistants promised by the end of the quarter.
  4. September 2026: Deutsch and Imbert publish the token framework, pairing cost metrics with value metrics and naming waste rather than use as the target of controls.

On SAP AI Core, generative use is metered in tokens that convert into capacity units, so even the vendor’s own platform speaks the language finance is being asked to learn. The invoice will not go away. What changes is the line it sits on. At SAP that line is already next to headcount, and only one of those two is still being asked to grow.

Harry is the editor and lead writer of NEWFOUND TIMES, an independent publication he owns and edits. He has ten years in journalism behind him, the first stretch as a reporter filing daily and the later ones running a desk, and he still reports most of what he publishes. Datasets are his preferred starting point: a spreadsheet from a statistics office, a results table, a public register, a sales report. He opens the data himself rather than relying on a summary of it, and every figure that ends up in an article is checked against that source. The site covers ten sections for readers spread across many countries, and business, science and technology sit next to news, sports, entertainment, lifestyle, travel, gaming and auto on the front page. Errors are corrected openly: the article is updated, the correction is dated, and the site's corrections policy explains how the process works. Readers can send data, documents or complaints to support@newfoundtimes.com and expect a reply from him.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending