JS
Back to All Articles

My AI Bill Went Up and My Cost Per Task Collapsed

August 4, 2026  ·  13 min read

My AI Bill Went Up and My Cost Per Task Collapsed

My AI bill has grown almost every month this year. I consider that evidence the strategy is working. If that sounds backwards, this article is for you, because the way most organizations read their AI spend is about to cost them the thing that was working.

This week I did something I ask my clients to do and most leaders never do. I read my own AI bill line by line, the way a CFO would read it if it were ten thousand times bigger.

Two numbers came out of that exercise, and they point in opposite directions.

The first: my monthly spend on AI infrastructure went from zero in February to roughly $250 by June. Up and to the right, month after month.

The second: between June and July, my cost per unit of work fell 20 percent.

Same system. Same month. Both numbers are true. And only one of them tells you whether the money is being spent well.

The invoice measures consumption. It does not measure judgment.

Here is the raw record from my own books, because I am not going to make an argument about AI economics using someone else's survey data when I have receipts.

  • January and February: $0. The system did not exist yet.
  • March: $13.73. First experiments.
  • April: $37.75. The first real workloads.
  • May: $159.44. The system took over scheduling, meeting prep, and follow-up.
  • June: $249.84. Full build-out.
  • July: $225.85 through the 30th.

Read that as a CFO reading a budget line, and you see a cost growing without a ceiling. Eighteen times bigger in four months. The instinct is to ask who approved this and where it stops.

Now add the second measurement.

AI usage is metered in tokens. A token is a fragment of a word; think of a million tokens as roughly a novel and a half of text processed. Tokens are the kilowatt-hours of this technology, and once you know your token volume, you can compute the number that matters: what a unit of work costs you.

In May, my systems processed about 113 million tokens. In June, 174 million. In July, 196 million.

So between June and July, the work done grew 13 percent. The bill went down almost 10 percent. The cost per million tokens dropped from $1.44 to $1.15.

More work. Less money. That combination is not luck. It is the signature of managed adoption.

The collapse had three causes, and none of them was waiting.

The unit cost did not fall because a vendor sent me a coupon. It fell because of three specific decisions, each of which has an equivalent at any scale.

First, the models themselves keep getting better per dollar. In late April I moved my workloads to a newer model generation mid-month. The work continued; the economics improved. This is the part of the curve you get for free, and it is real. It is also the smallest part.

Second, right-sizing. In mid-July I audited every AI workload in my operation and found three tasks, all simple classification, running on a model far more capable than the job required. Sorting incoming email into categories does not need the senior analyst. It needs the intern with a checklist. I moved those three tasks to a smaller, cheaper model, and you can see the exact week it happened in my usage data.

Third, engineering. A week later I shipped prompt caching, which is a plain idea with an unglamorous name: stop paying full price to re-send the assistant the same background information on every request. Send it once, reference it cheaply afterward.

Model improvements, right-sizing, caching. One arrived from the outside. Two were management decisions somebody had to own and make.

Unit costs do not fall on their own. They fall because someone is responsible for making them fall.

There is one more piece of my cost structure worth naming, because it applies directly at enterprise scale: the heavy engineering work, the building and rebuilding of the system itself, runs on a flat monthly subscription. Two hundred dollars, fixed, no meter. The variable spend covers production workloads; the unbounded work lives under a price ceiling. Two-part cost structures like this are how you keep an AI budget from becoming a fear.

Your numbers have more zeros. The shape is the same.

I run a consulting practice, deliberately lean, and my entire AI operating system, the one that schedules my meetings, drafts my correspondence, prepares my briefings, and reconciles my books, has cost me less than $700 in metered usage this calendar year.

I am not publishing that number to impress anyone. I am publishing it because I ask clients to instrument their AI costs, and I am not willing to ask for transparency I do not practice.

At your scale, the same curve carries six or seven figures. The mechanics do not change. Work volume grows as adoption succeeds. Unit rates fall as models improve and as someone does the unglamorous work of routing, caching, and right-sizing. The two forces run against each other, and the invoice you see each month is the net of that collision.

Which means the invoice, by itself, is unreadable. A rising bill can mean runaway waste or compounding adoption. A flat bill can mean discipline or abandonment. You cannot tell the difference without the unit number.

The mistake is about to be made in a budget meeting near you.

Here is why I wrote this piece now.

AI spend is moving out of the experimentation budget and into the operating budget, which means it is about to get the standard operating-budget treatment: year-over-year comparison, variance analysis, and pressure on any line that grows.

And a line that grows is exactly what successful AI adoption looks like.

The organizations getting real advantage from this technology will see their AI bills climb for years, because the systems keep absorbing more work. If finance reads that growth the way it reads most cost growth, the response will be caps and freezes. The caps will land on the systems doing the most work, because those are the ones that cost the most. The result is an organization that punishes its own adoption curve.

The fix is not a bigger budget. The fix is a different denominator.

A cost per task. Cost per claim processed. Cost per ticket resolved. Cost per contract reviewed, per candidate screened, per report drafted. Whatever the work is in your organization, that is the unit, and the question for every AI line item becomes the question my July numbers answer: is the unit cost falling while the volume grows?

When the answer is yes, the rising bill is an asset compounding. When nobody can answer, the bill is unmanaged, whatever its size.

There is a second budgeting error hiding in the same meeting: anchoring multi-year plans to today's rates. My cost per million tokens fell 20 percent in a single month, and a meaningful share of that came from the technology itself improving underneath me. Pricing a three-year AI initiative at today's unit rates is how projects get declined that should have been approved.

What I would take into your next budget conversation.

Three moves, all available regardless of scale.

  1. Demand the unit, not the total. Every AI initiative in the portfolio should report cost per unit of work next to total spend. If a team cannot name its unit, that is the finding.
  2. Give every AI workload an owner who revisits the model choice on a schedule. The July audit that cut my classification costs took one session. It sits on my calendar quarterly now. At enterprise scale that review is worth someone's whole job.
  3. Treat optimization as an operating practice, not a project. Caching, routing, right-sizing. None of it is exotic. All of it is the difference between the $1.44 and the $1.15, and the gap widens every quarter it goes unowned.

Then set the expectation out loud with your finance leadership: the total will rise, the unit will fall, and you intend both, on purpose. Write it into how the line is reviewed before the first variance report arrives, because retrofitting that understanding after a budget fight is far harder.

The meter on the wall tells you how much intelligence you bought.

Only the unit cost tells you whether you knew what you were doing with it.

My bill went up. My cost per task collapsed. One of those numbers is the strategy.

Ready to scale with clarity?

I take on 3–4 new clients per quarter.