What it costs to run AI features in a real product
An AI feature costs money every time someone uses it. The provider charges for the text you send to the model and for the text the model writes back, and larger models charge more for the same amount of text. So the bill grows with usage, which is what catches people out. Most of that cost is under your control, though. You choose the right size of model for each job, stop paying full price to send the same text again and again, delay work that does not need to happen straight away, and set hard limits so a bad day cannot turn into a bad month.
Below we explain what you are paying for, why the numbers surprise people, and what we do in FFSN, the fantasy football news product we build and run ourselves.
What you are paying for
AI providers such as Anthropic, which makes the Claude models, charge per token. A token is a small piece of text, roughly a piece of a word. You pay for tokens in both directions.
The first direction is input: everything you send to the model. That includes what the user typed, plus the instructions, examples and background information your product adds to every request behind the scenes. All of this together is called the prompt. The second direction is output: everything the model writes back.
Longer prompts cost more. Longer answers cost more. And the model you pick sets the price of each token. Providers offer models in several sizes. The larger ones write and reason better and cost more per token. The smaller ones cost less and are perfectly capable of simple jobs.
So the cost of one request comes down to how much text goes in, how much comes out, and which model handles it. Multiply that by how many requests your product makes in a day and you have the bill. A feature that writes a long article will cost far more per use than one that sorts a support email into “billing” or “technical”, because it produces much more text and probably needs a larger model to do it well.
We are not quoting prices here on purpose. They change often, so any figure we printed would be out of date before long. Check the provider’s pricing page when you plan a feature.
Why AI bills surprise people
Most running costs of a website are roughly flat. Hosting and a domain cost about the same whether ten people visit or ten thousand. An AI feature behaves more like a utility meter. If usage grows tenfold and nothing in the design changes, the AI part of the bill grows roughly tenfold with it.
The second surprise is the hidden part of the prompt. A user might type one sentence, but the product sends pages of instructions and context along with it, on every single call. That hidden text is often most of what you pay for on the input side.
The third is calls nobody counted. A feature that runs on every page view, a check that runs on every event, an automatic retry after an error, or a bug that sends the same request in a loop. Each one is a paid request.
The last is using the biggest model for everything. It is the easy default during development, and also the most expensive option. Many tasks do not need it.
Spending less on each call
FFSN uses Claude models to write fantasy football news and analysis. You can read more about how it works on the FFSN project page. These are the techniques in its code that bring down the cost of each request.
Matching the model to the job. Long articles are written on the most capable model, because quality matters most there. Small jobs, such as short interview messages and simple checks, run on a smaller, cheaper model. Paying top rates for a short message adds up quickly across many calls.
Prompt caching. The provider can remember the opening part of a prompt it has seen recently and charges less when the same opening arrives again. To take advantage of that, we lay out every prompt with the large, unchanging part first (the standing instructions and background) and the part that changes at the end. The writing call, the editing call and the reply call all share that same opening, so the most repetitive text is billed at the reduced rate.
Batching work that can wait. Articles that are scheduled in advance do not need to be written the moment we ask. We send them through Anthropic’s batch service, which accepts a group of requests and returns the results later. It costs half as much as an immediate request, in exchange for the wait. If a batch does not come back in time, the system falls back to writing that article directly at the normal price, so the article still goes out.
Making fewer calls
The cheapest AI call is the one you do not make.
Write once, reuse many times. Some news affects every league at once, such as an NFL injury. For that kind of story, FFSN has the AI write one take. The short note that makes it relevant to a particular league is filled in by ordinary code, with no AI call at all.
Doing the cheap thing first. Plenty of events are not worth an article. FFSN decides which ones are with ordinary code, which costs next to nothing to run. The AI only gets involved once there is something worth writing about.
Both ideas apply well beyond sports news. Before any step calls a model, it is worth asking whether plain code could make the decision, or whether one result could serve many customers.
Keeping the bill predictable
A daily spending cap. One FFSN feature has a daily limit on AI spending, so a bug or a busy news day cannot run up an unlimited bill. If you add a cap, decide in advance what the feature does when it is reached, because something has to give.
Recording every call. Each AI call records how much it used and which feature made it. That lets us see what each feature costs and spot one that starts to creep before it shows up as a surprise on the invoice.
Rate limits per customer and per feature. Limits on how often each customer and each feature can call the AI stop one very heavy user, or one misbehaving feature, from driving the whole bill.
Switching models without a release. Which model each feature uses is stored as a setting, outside the code. If a model starts behaving badly, or prices change, we can move a feature to a different model without shipping a new version of the product.
What to ask before approving an AI feature
If someone is building an AI feature for you, these questions will tell you quickly whether the cost has been thought through:
- What happens to the bill if usage grows tenfold? You want a rough answer, and to know which part of the feature grows fastest.
- Is there a spending cap, and what does the feature do when it hits it?
- Which step is the expensive one, and can it be batched or cached?
- Can we see the cost of each feature separately?
- Does every step need AI, or could ordinary code handle some of them?
- If a model gets more expensive or starts misbehaving, can we switch it without a new release?
Good answers do not need to be precise. A developer who has not considered these questions at all is the warning sign.
Where to start
You can test the most important question yourself before spending anything on development: paste a few real examples of the task into a chat tool such as Claude and see whether the results are good enough. If a smaller, cheaper model handles it well, that is a good sign for your running costs.
If the idea holds up and you want it built into your product with these controls in place from the start, that is what our AI integration service covers. You are welcome to get in touch to talk it through.