The number on the slide was correct. That was the problem. Four hundred and twelve million tokens, forty-seven thousand nine hundred dollars, broken out by model and by week, and the finance lead asked what it bought. Not rhetorically. She wanted the sentence. And nobody in the room could produce one, because there is no sentence that starts with a token count and ends in something a business recognizes as a purchase.
I have sat in some version of that meeting four times this year, and the shape is always the same. The marketing side is not hiding anything. The invoice is itemized, the attribution by team is clean, and the finance side has done nothing wrong either. Everyone is looking at a true number that cannot answer the question being asked of it, and the meeting resolves the way those meetings do, with somebody telling a story about a campaign that went well.
Nobody is buying tokens. A token is an input meter. It got promoted to a budget line during a stretch when nobody was checking, and now that everyone is checking, the meter cannot answer. This is not a complaint about vendor pricing and it is not the argument that AI is overpriced. It is narrower and more irritating than that. The unit is wrong. A budget line denominated in the wrong unit cannot be defended, cut, or renewed. It can only be argued about, which is exactly what this industry has spent the year doing.
A token is not a unit of cost
Start with the simplest version. As I write this in August 2026, a million input tokens on a small model runs about twenty cents. A million output tokens on a frontier model runs about twenty-five dollars. Same count. More than a hundred times the money. Turn on prompt caching and the cheap end drops another ninety percent; run it through a batch endpoint and everything halves again. The spread between the cheapest and most expensive million tokens you can buy is now wide enough that the count tells you almost nothing about the cost.
So when a budget is denominated in tokens, you are budgeting in keystrokes. Keystrokes are real. They are measurable, they correlate loosely with effort, and a report of them would be entirely accurate. No finance function on earth would accept one.
And in an agentic system it is not a unit of work either
Here is the part that took me a while to see, and it is the reason this is a strategy problem rather than a procurement problem.
In a single-shot tool, token spend roughly tracks how much you asked for. In an agentic system it tracks how hard the request turned out to be. A retry loop, a reflection pass, a tool call that fails and forces a re-plan, a context window read four times because the first three passes did not resolve: all of that is the system working harder on a harder problem, and all of it lands on the invoice looking identical to the system working on more problems.
Which means an AI bill that doubles is genuinely ambiguous between two facts. Either you bought twice as many decisions, or you bought the same decisions from a system that was twice as unsure. Those call for opposite responses. The first is a scaling story and you should probably fund it. The second is a quality story and you should probably fix it before you fund anything. No cost attribution scheme in existence separates them, because both appear as consumption, and consumption is the only thing being counted.
Set that next to the numbers everyone is quoting. Average enterprise AI budgets have gone from roughly one point two million dollars in 2024 to about seven million in 2026. Per-developer token consumption is up something like eighteen times in nine months. Seventy-three percent of enterprises say their AI costs came in over projection. Every one of those is being read as a cost-control failure. I think it is mostly a units failure wearing a cost-control costume. You cannot forecast a quantity when you do not know what drives it, and difficulty is not a thing anybody forecasts.
The argument nobody can win
Which brings us to the loudest fight in enterprise software right now.
MIT reported that ninety-five percent of enterprise generative AI pilots produced no measurable impact on the P&L, based on fifty-two executive interviews, a hundred and fifty-three leader surveys, and three hundred public deployments. The critics have a fair objection: the study measured rapid P&L impact within about six months, and it does not capture efficiency gains, churn reduction, or pipeline velocity. Both sides have been at this for months.
Neither side can win, and it has nothing to do with the quality of either argument. Proof is a ratio, and this one has no denominator. Ninety-five percent of what, per what, is a question the industry has not answered, so "did the pilot work" is not yet a well-formed question. Two competent parties are arguing about a fraction with an empty bottom half.
The same hole shows up in the survey data and it is more uncomfortable there. Ninety-six percent of CMOs report that AI is driving end-to-end transformation of their function. Roughly a third have changed anything about how their marketing systems actually work, and BCG's June read had eight percent running campaigns where multiple agents operate autonomously. I do not think that gap is dishonesty. "Is AI transforming your function" is a question with no unit in it, so it collects an impression rather than a measurement, and the impression is what travels upward into the board deck. Ask instead how many decisions of which class the org delegated last quarter and what each one cost, and the eight percent and the ninety-six percent stop contradicting each other, because only one of them was ever a number.
A dollar-denominated token budget is a better answer to the wrong question.
Two people who will push back, and both of them are right
The first is the FinOps lead, and she gets there before I finish the sentence. They do not budget in tokens. They budget in dollars, with attribution by team and workload, chargeback, the whole apparatus. This is true and it is genuinely ahead of where most marketing organizations are; if your AI spend is dollar-denominated with real attribution, you are running a tighter shop than the median. But attribution tells you which team spent it and chargeback tells you who pays for it. Neither says what was bought. And what was bought is the only question a renewal meeting is actually asking, which is why those meetings still end in anecdote even in companies with excellent cost allocation.
The second is the platform engineer, and his objection is stronger. Token cost per interaction is load-bearing for him. It is how he picks a model, sizes a context window, decides what to cache, and catches a runaway loop at two in the morning. He is right, he should keep it, and nothing here asks him to stop. The failure is not that anyone measures tokens. It is that the token number is the only number that survives the trip from his dashboard to the budget conversation. An operational metric got promoted into a financial one somewhere in the hallway, and nobody noticed because it was the only number available.
The invoice is the new thing
This site has an anchor essay arguing that every marketing KPI measures activity rather than decisions. Fair to ask whether this is that essay with a different noun. It is not, and the difference is worth being precise about, because it is the whole argument.
Cost Per Decision prices the decisions you make. This prices the decisions you have handed off. And handing off is what changed, because when a human made the call, the cost of deciding was salary. It was buried in headcount, spread across a dozen cost centers, and nobody ever produced a line item for judgment. Delegate the same call to a system and the cost of deciding becomes a discrete charge on a supplier invoice, arriving monthly, denominated in a unit the supplier picked.
That last clause should be familiar if you have been reading along. Whoever defines the unit defines the denominator, and a number you did not construct is a quote rather than a measurement. This site has spent two months making that case about ad platforms and never once turned it on its own vendor bill. The token is a seller's unit. It is a perfectly honest one, chosen for good engineering reasons rather than commercial ones, which somehow makes it more slippery: nobody is being deceived, so nobody thinks to check.
What goes on the line instead
McKinsey has arrived in the neighborhood and deserves the credit. Their framing is to move off token counts toward business-outcome metrics: cost per claim processed, cost per resolved ticket, cost per completed workflow. That is a real improvement and any team that adopts it is better off.
It also stops working exactly where marketing starts. Those are task units, enumerated one industry at a time, and they need the work to be countable and repetitive. Take an agent that spent a quarter reallocating budget across fourteen campaigns. How many workflows was that? The question does not have an answer. There is no natural unit of throughput, because the output was not a processed item, it was a sequence of judgments that foreclosed alternatives you will never see.
The general case needs three things and none of them come from the vendor. Name the decision class: budget reallocation across a defined portfolio, creative rotation within a fixed set, bid tolerance inside a stated band. State the rate: how often a decision in that class actually gets made. Name the owner, meaning the person still accountable for the outcome, because an agent is not an owner and an unowned decision has no denominator. Divide the spend by the count and you have a cost per decision. It will be a rough number the first time. It will still be the first number in the conversation that can be compared to something: to what those decisions cost when people made them, and to what they are worth when they go right.
None of this is new to anyone who runs paid media, incidentally. Performance Max meters impressions and clicks while what you are actually buying is a sequence of allocation decisions across five signal layers, and those two objects have never been the same thing. The ad platforms have been billing for inputs and delivering decisions for years. It is the oldest instance of the pattern, which is presumably why the industry that lives inside it did not recognize the shape when it showed up on a different invoice.
Sometimes a decision class costs more through the machine and you run it anyway, for speed or coverage or because the alternative is not staffing it. Legitimate call. It is also a priced exception rather than a rounding error, and there is a calculator here for writing down what it was supposed to buy. The point of a number is never that it flatters you. It is that a decision class you have priced can carry a kill condition, and one you have not priced cannot, because there is nothing to compare the threshold to.
This site argues more than it predicts, and an argument cannot be scored in retrospect. So here is one that can be, with a date on it.
On February 1, 2027, come back and check this. The first durable pricing standard for agentic marketing will be denominated in decisions or workflows rather than tokens, and it will arrive because a vendor started printing that number to win deals, not because buyers demanded it. Buyers do not demand a unit they cannot name.
If on that date the major agentic marketing platforms are still quoting per token or per seat, and no buyer group has published a decision-unit standard, I was wrong about the mechanism. I would rather know which way, and why, than be vague enough to have been neither.
The meeting you are going to have in October
Budget season is close, and your version of that meeting this autumn will be worse than the one I opened with, because the number will be bigger and the patience thinner. Somebody will ask what the AI spend bought. If the only answer is a consumption figure, the spend gets defended with a story, and spend defended with a story gets cut by whoever holds the spreadsheet.
The work to avoid that is unglamorous and takes an afternoon. Pick the two or three decision classes you have actually delegated. Count how many decisions of each got made last quarter. Divide. You will not like the first number, and it will be wrong, and neither matters, because a rough figure in the right unit beats a precise one in the wrong unit in every conversation that decides anything.
The token number is not going away and it should not. It is the right instrument for the people tuning the machine, and they need it. It simply cannot travel. What travels into the room where budgets live has to be denominated in something the business already recognizes as a purchase, and the only candidate on the table is the decision itself. You cannot renew a budget line you cannot describe. Write down what one delegated decision is worth to you before the invoice teaches you what it costs.