Cost per task is the new AI price tag.

Owners keep asking which model is cheapest. It is the wrong question. The model bill is usually the smallest line on the invoice. The number that decides whether AI pays for itself is what one finished task costs, including the minutes a person spends fixing it.

Jastej Singh Sehra Jastej Singh Sehra · LoopSuit, Vancouver 8 min read
The short answerCost per task = (model price × tokens the task uses × attempts to get it right) + (share of tasks a person has to fix × minutes to fix × their hourly cost). For most business workflows the model part is a few cents; the fix part is dollars. So pick the model per step by measuring failure rates, not by comparing per-token prices, and expect a different answer where staff time costs less.

Every few weeks a new model launches and my inbox fills with the same question in different words: "Should we switch? It's cheaper." The comparison is always per-token price, because that is the number on the pricing page.

Per-token price is real, but it answers a question nobody running a business is actually asking. You don't buy tokens. You buy finished work: an enquiry answered, an invoice read, a quote drafted. And the cost of finished work has three parts, only one of which is on the pricing page.

$ per million tokens
Model priceper token
×
Tokensper attempt
×
Attemptsto get it right
+
Human fixesshare × minutes × wage
= cost per finished task
Figure 1. The pricing page only shows the first box. The last box, the minutes a person spends correcting the AI, is usually the biggest.

What the models cost right now

For reference, here is Anthropic's current API pricing for the models we use most, per million tokens (source). Roughly, a million tokens is 700,000 words of reading or writing.

ModelInputOutputWhere we use it
Claude Fable 5.1$10$50Hard reasoning, long multi-step work, the "judge" that checks other steps
Claude Opus 5.5$4$20Drafting that customers read, complex extraction, most agent steps
Claude Sonnet 5$2$10Everyday drafting, classification with nuance, summaries
Claude Haiku 4.5$1$5High-volume sorting, tagging, pulling fields out of clean inputs

Other providers sit in similar bands, and open-weight models keep pushing the floor down. The spread from cheapest to most capable is about ten times. That sounds like it should decide everything. In practice it almost never does, and the next section is why.

Why the model bill is the smallest line

Take a real shape of task: a property manager's AI reads a tenant email plus the lease, and drafts a reply. That's roughly 20,000 tokens in and 3,000 out. On the most expensive model above, one attempt costs about 37 cents. On the cheapest, under 4 cents.

Now suppose the cheap model gets one in eight replies wrong enough that someone has to rewrite it, and each rewrite takes six minutes of a coordinator's time. At a Vancouver office wage of around US$40 an hour, that one fix costs about $4: the same as running the cheap model a hundred times, or the most expensive one about ten times. Tokens cost cents. Mistakes cost minutes, and minutes cost dollars.

Cheapest model3.7¢ a try · 12% need a fix$0
Mid-priced model15¢ a try · 3% need a fix$0
Most capable model37¢ a try · 2% need a fix$0
Model billStaff time fixing mistakes1,000 tasks a month · US$40/h
Figure 2. Monthly cost of 1,000 tenant replies (Haiku 4.5, Opus 5.5 and Fable 5.1; failed tasks get a second model pass). The cheapest model looks cheapest until you count the fixes; the most capable model is overkill for this job. Illustrative failure rates; model prices are Anthropic list prices.

Notice that the winner is neither end. That is the pattern we see on most client workflows: the right model is the cheapest one whose failure rate is already low on your tasks, and you only find it by measuring.

Work out your own cost per task

Play with the numbers. The bit most people underestimate is the fix rate: count a task as failed if a person had to change anything a customer would notice.

Who fixes mistakes?

Fix rates per model are illustrative starting points: 12% · 5% · 3% · 2%. Replace them with numbers from your own test set. Prices assume 85% of tokens are input, at Anthropic list prices, no caching or batch discount.

Figure 3. Switch the office. When the person fixing mistakes costs less, cheaper models become competitive again. That is a real design decision for teams split between Canada and India.

The Vancouver versus Delhi switch is not a gimmick. We work with teams in both places, and the same workflow genuinely deserves a different model mix depending on who reviews the output. With a Canadian coordinator reviewing, you pay for a stronger model and fewer mistakes. With a reviewer in Delhi or Gurgaon, a cheaper model with a slightly higher fix rate can be the smarter build, as long as the mistakes are the catchable kind.

The cheapest model is rarely the cheapest system. The most expensive model is rarely needed.

Choose the model per step, not per company

"Which model do we use?" is the wrong unit of decision. A good workflow has several steps, and each step deserves its own answer. Reading and tagging an email is easy; writing the reply a customer reads is harder; checking a refund against policy is where you want the strongest reasoning.

One workflow, three models

The owner's monthly view. Each step runs on the cheapest model that passes its own quality bar, and the bill is shown per finished task, not per token.

  1. The month at a glance: cost per finished task, not tokens.
  2. Open one workflow to see what each step runs on.
  3. Sorting runs on the small model; replies on the mid model; policy checks on the strongest.
  4. A suggestion: one step can safely move down a tier.
Figure 4. Cost reported per finished task, with each step on its own model. Illustrative business and numbers.

The method we use is top-down. Get a step working on the most capable model first, so you know what "good" looks like. Then try the next tier down against a fixed test set of real examples, and keep stepping down while the quality holds. Stop the moment it doesn't. Going bottom-up ("start cheap, upgrade if it breaks") sounds thrifty, but you never find out how good the step could have been, and you debug prompts that were never the problem.

This only works if the test set exists. A few hundred real, anonymised examples with the right answer marked is the most valuable asset in any AI system we hand over. It is what lets a client change models next year without guessing.

The four levers that actually cut the bill

Before you swap models, pull the free levers first. In order:

0%
Caching

Instructions, policies and documents that repeat on every request can be cached. Anthropic lists cached reads at up to 90% off.

0%
Batching

Anything that can wait minutes (nightly summaries, bulk tagging) can run through the batch API at half price.

0–60%
Trimming

Send the model the paragraph that matters, not the whole thread. Typical savings we see from trimming inputs. Illustrative.

The fourth lever is the one in the calculator: lower the fix rate. A better prompt, a clearer approval screen, or a validated input field (a date picker instead of free text) often saves more than any model switch. Fewer fixes is also the only lever that improves the customer's experience at the same time.

A note on news cycles. New models will keep arriving every few months, each cheaper per unit of intelligence than the last. If your system is built with a test set and per-step model choice, a new launch is a one-afternoon experiment. If it isn't, every launch is a rebuild. Design for the switch.

How we build this at LoopSuit

On every system we ship, cost is reported the way the owner thinks about it: per finished task, per workflow, per month, in a screen like the one above rather than a usage export. We set each step's model top-down against a test set we build with the client, cache and batch by default, and hand over the test set so the model choice stays theirs.

If you are weighing an AI project and the maths feels hand-wavy, that is the first thing we fix. It is usually a half-day exercise, and it often ends with "automate less than you planned, and automate it properly". See the AI Clarity Session, or read why the approval screen is where fix rates are won or lost.

Questions people ask us

How much does it cost to run an AI automation per month?

For a typical small-business workflow (a few thousand emails, enquiries or documents a month), the model bill is usually tens of dollars, not thousands. The bigger costs are hosting, monitoring, and the staff time spent correcting outputs. Work it out per completed task, then multiply by volume.

Is the cheapest AI model the cheapest option?

Rarely. A cheaper model that gets the task wrong more often creates more human fixes, and staff minutes cost far more than tokens. In most workflows we measure, a mid-priced model wins on cost per completed task. The only way to know is to test on your own tasks.

What is cost per task in AI?

It is the all-in cost of getting one unit of work finished correctly: the tokens the model uses, any retries, and the human time to review or fix the result. Some model providers now publish cost-per-task benchmarks for the same reason: per-token prices alone don't predict the bill.

Does it cost the same to run AI in India and Canada?

The model price is the same everywhere (API usage is billed in US dollars). The difference is the cost of the human fixing mistakes. Where staff time is cheaper, a cheaper model with a higher error rate can still be the better choice. Where it is expensive, paying for a stronger model usually pays off.

How can we lower our AI running costs?

In order: cache the context that repeats on every request (Anthropic lists up to 90% off cached reads), batch anything that doesn't need an instant answer (50% off), trim what you send the model, then step down to cheaper models only for the steps where quality holds on your own test set.

Want your real cost per task?

Bring one workflow and last month's volume. We'll estimate the model bill, the fix rate and the staff time, and tell you honestly whether it's worth automating.