Summary: Agent pricing in 2026 is four models wearing a dozen names: per-seat, per-token, per-task, and per-outcome. Each one allocates a different risk between the vendor and you. This guide explains what each model optimizes for, the failure mode of each, and the four questions that turn any pricing page into a comparable cost per decision.
Per-seat: the SaaS reflex
Copilots and assistant products charge per user per month, because that is how software has been sold for twenty years. It works when the product amplifies a human at their desk. It breaks the moment the point is autonomy: an agent that processes 10,000 cases a month is not a seat, and seat counts do not move when the work does. Enterprises that buy agents on seats routinely discover they are paying for licenses while the actual unit of value, completed work, goes unmetered and unmanaged.
Bites when: usage explodes per seat (vendor’s margin problem becomes a price hike) or automation replaces the very seats being counted.
Per-token and credits: the API reflex
Consumption pricing passes model economics straight through: tokens, or abstracted credits that map to tokens and compute-seconds. It is honest in one sense (you pay for what runs) and treacherous in another: you pay for how the vendor’s architecture runs, not for what you get. A measured study found growing-context agents replay 3.6x the input tokens of a single compiled pass for the same verdicts; under consumption pricing, that multiple is your bill. Retry storms, verbose reasoning, and voting ensembles all monetize as your cost.
Bites when: the architecture is inefficient, workloads spike, or a workflow change triples consumption overnight. Budget variance is the defining complaint; our token economics guide covers the forecasting math.
Per-task: metering the work
Task pricing meters discrete units of agent work: a document processed, a policy evaluated, a governed operation executed. Done well, it aligns three things at once: your cost scales with work completed, the vendor is incentivized to execute efficiently (their margin improves as their architecture improves, instead of billing you for waste), and finance can forecast by multiplying volume by rate. This is the model MightyBot uses, with tiered rates that fall as monthly task volume grows, and it is why the ROI calculator can quote a 3-year platform TCO in four inputs.
Bites when: the task unit is vague. Demand a written definition of the unit and a worked example: how many tasks is one of your typical cases?
Per-outcome: pricing the result
The frontier model: price per completed business outcome, a resolved claim, a funded loan review, sometimes with quality terms attached. Maximum alignment, hardest to contract: outcome attribution, edge-case ownership, and quality disputes all need language. Expect this model to expand as accuracy becomes contractable; systems with why-trails have an advantage here, because provable decisions are billable decisions.
The four questions that normalize any pricing page
- What does one completed decision cost at my volume? If the vendor cannot answer in dollars, you are buying a meter.
- What is the variance? Ask for the P90 case cost, not the average. Consumption models hide their risk in the tail.
- What does failure cost? Retries, timeouts, and human escalations: which of these tick the meter?
- What happens at 10x volume? Tiers should fall with scale; meters that stay linear are margin, not cost.
Then put every vendor’s answers into the same arithmetic: annual volume times cost per decision, plus platform and implementation fees, over three years. That is the comparison the pricing pages are designed to prevent, and the one the build-versus-buy analysis walks end to end.