FREE TOOL

Updated July 2026

Free AI Agent ROI & TCO Calculator

Estimate the real 3-year TCO of building production AI agents before you commit the team.

Model workflow volume, human labor, engineering FTE, launch delay, token usage, infrastructure, and residual review. Then compare the internal build path against MightyBot's platform + FDE path.

Default build scenario

$7.5M estimated internal 3-year build TCO
$6.8M estimated savings with MightyBot path

What the MightyBot ROI/TCO Calculator measures

The MightyBot ROI/TCO Calculator estimates whether an AI-agent automation program creates a positive business case after the full production cost is counted. It compares the human-only baseline against internal build cost, model and token usage, infrastructure, governance, residual review, accuracy, time to production, and the MightyBot platform + FDE path over a multi-year period. The goal is a defensible build-vs-buy plan, not a rough automation estimate.

The short answer: an AI agent TCO calculator should count the full cost to build, govern, run, monitor, and review production agents, then compare that total against the measured human-only workflow baseline and the platform alternative.

Best measurement method Use measured workflow volume, human handling time, loaded labor cost, launch delay, residual review, and full production TCO.
ROI proof Show the baseline, the automated operating model, the payback period, and the assumptions behind each cost and savings line.
Financial services Include compliance review, audit evidence, access control, exception handling, and human approval gates in the ROI model.
ROI formula: (human-only baseline cost - automation TCO) / automation TCO. TCO formula: implementation + platform or engineering + model usage + infrastructure + governance + residual review + maintenance. Build-vs-buy test: compare 3-year cost, launch delay, accuracy, governance, and operating ownership.

Formula

How to calculate AI agent ROI

AI agent ROI = (human-only baseline cost - automation TCO) / automation TCO. This calculator treats TCO as a 3-year model, not just the first platform invoice. It also supports build vs buy AI agents planning by comparing AI agent implementation cost, the cost to build AI agent infrastructure internally, and platform TCO.

1. Set the human baseline Workflow volume, human manual labor minutes, and loaded reviewer cost.
2. Model AI agent implementation cost Engineering labor, launch delay, architecture, tokens, infrastructure, governance, QA, and the cost to build AI agent systems internally.
3. Compare 3-year TCO Human-only operations, internal build, and MightyBot platform plus FDE.

Calculator

Calculate AI agent ROI and 3-year TCO

No sign-up required. Results update as you change assumptions.

Agent Architecture and LLM token usage
Single document or record per case with a handful of checks. Roughly 5 billable tasks per workflow in production.
Common prototype pattern with repeated context replay, retries, and validation loops.
Architecture default. Includes extraction, retrieval, policy checks, validation, and exception handling.
Architecture default. Includes prompt, context, retrieved evidence, tool results, and output.
Input tokens dominate agent workloads and are the part prompt caching can reduce.
Measured output share on document evaluation workloads is roughly 9% of total tokens.
50%
Share of input tokens served from prompt cache. Output tokens are never cached.
Internal build team
Senior engineers, architects, data/ML, and platform support.
Salary, benefits, taxes, recruiting, equipment, and management load in U.S. dollars.
Months before the first regulated workflow is production-ready.
FTE for evals, monitoring, integrations, model upgrades, and support.
Voting trades tokens for accuracy: it lifts effective accuracy but floors near a 10% borderline error rate and multiplies model cost.
Workflow volume and human capital costs
Cases, documents, images, reviews, or decisions per month.
Minutes of human review, validation, and handoff.
Used to estimate the manual-work baseline in U.S. dollars.
Elapsed days from intake to decision, including queue time.
Interest carry, SLA penalties, or working capital per case-day. $0 keeps the model cost-only.
Internal infrastructure costs
Auto-estimated one-time setup for data prep, IAM/security review, audit logging, observability, compliance, and integrations.
Auto-estimated annual non-model Opex for compute, memory/state stores, observability, eval tooling, queues, storage, and support systems.
Infrastructure model At 25,000 workflows/month with 8 model-backed steps per workflow, this model estimates $100K one-time setup and $100K/year for non-model infrastructure and controls. Model API token spend is calculated separately.

Methodology

What this AI agent TCO model includes

The calculator separates human-only operating cost, prototype cost, and production cost. It includes human labor baseline, quality-adjusted residual review effort, engineering capacity, setup work, ongoing platform operations, model usage, retry behavior, annual model-efficiency improvement, and the cost of running a less efficient agent architecture at production volume. Use it as an AI agent business case model when deciding whether to build internally, hire an integrator, or buy a production agent platform. If you are deciding whether to build or buy AI agent platform capabilities, this section shows which cost lines belong in the business case.

(MightyBot path: use the platform with your team or MightyBot FDE, including task-based platform pricing and implementation assumptions.)

Model efficiency assumption: internal token costs improve 25% in year two and again in year three; MightyBot platform planning starts from the selected complexity band (5 billable tasks per workflow for simple operations) and applies the same annual efficiency curve over the same period.

Workload complexity scales both sides symmetrically: heavier cases raise the internal build's token volume and the MightyBot billable-task rate together, so the comparison holds whether a workflow consumes 5 tasks or 60. Production workflows can get chunky: the heaviest compliance agents, such as covenant monitoring and SAR drafting, average 100 to 200 tasks per workflow because every task is a discrete governed operation with its own audit trail.

Speed is priced, not just described: the optional speed dividend values faster turnaround per case-day, and it accrues only to cases that clear without returning to the human review queue, from the month each path actually ships.

Internal build + tokens Engineering FTE, timeline, model usage, retry behavior, and token spend at production volume.
Internal infrastructure One-time setup plus annual cloud, tooling, monitoring, evals, queues, storage, and compliance Opex.
Human labor ROI Human-only baseline, residual QA/rework effort, and ROI against manual operations.
Prompt caching Separate input and output token prices with a cache-attainment assumption: cached input reads bill near 0.1x, while output tokens are never cached.
Quality equalization Close the build's accuracy gap with residual human review, or with self-consistency voting that multiplies model spend and floors near a 10% borderline error rate.
Speed economics Optional value per case-day of faster turnaround. The dividend accrues only to cases that clear without returning to the human review queue, from the month each path goes live.

Architecture assumptions

How agent architecture changes AI agent TCO

Architecture determines the cost curve. This dropdown models the current or in-house solution you are comparing against MightyBot. Sequential agents, RPA plus LLM steps, custom multi-agent systems, and deterministic workflows have different token, retry, observability, and maintenance profiles.

Two measured effects matter most. First, prompt caching compresses input costs, but voting strategies multiply output tokens, which caching never touches. Second, buying reliability with extra passes converges to an error floor, not to zero: a measured 20% borderline flip rate drops to roughly 10% at 3-vote majority while model spend more than triples. Both effects come from a July 2026 measured study on a reproduced multi-document evaluation workload.

ReAct / sequential prompt-chain agent Common prototype pattern with repeated context replay, retries, and validation loops.
RPA plus LLM workflow Automation workflow with LLM steps layered onto recipes, bots, or task flows.
Custom multi-agent framework Engineer-built orchestration with agent roles, tool use, memory, and observability. Measured growing-context agents replay 3.6x the input tokens of a single pass.
Deterministic workflow plus LLM calls Stronger internal-build option with explicit logic, targeted model calls, and production operating controls.

Speed economics

Speed is ROI: what each path costs you in time

Cost is only part of the business case. The other axis is time: how fast each case clears, and how fast the system reaches accuracy you can trust. Every point below production-grade accuracy is a case sitting in a human review queue, and queue time is money. In lending, a review that clears in one day instead of three removes two days of interest carry on the full loan value. In insurance and other SLA-bound operations, faster review is the difference between penalty and margin. Use the optional speed inputs in the calculator to price that dividend for your workflow.

100% 90% 80% 70% 60% 50% 0 12 24 36 Months from project start Human baseline 80% Your review queue lives in this gap
Modeled from your current calculator inputs. Both paths start from the 80% human baseline. MightyBot reaches 99%+ at go-live; the build path launches at prototype accuracy and climbs toward a measured floor. Both sides receive the same 25% annual efficiency credit. The difference is where they start: platform accuracy is production-proven at go-live, while build accuracy is a projection of future engineering effort.
Time to 99%+ production accuracy MightyBot: month 1, at go-live Internal build: not within 3 years
Modeled cycle time per case MightyBot: 2.1 min/case Internal build: 13.9 min/case
Cases still in the human review queue (first production year) MightyBot: 2% of cases Internal build: 64% of cases

Citations

AI agent TCO assumptions and cost evidence

The calculator defaults are planning assumptions, not a claim that every workflow behaves exactly this way. They are grounded in published agent architecture research, production framework documentation, independent AI cost analysis, and a first-party measured study (July 2026) that ran single-pass, agentic, and voting configurations of the same multi-document evaluation workload on live models and counted every token from API usage records. Task-per-workflow bands come from 2026 production billing observations across deployed workflows. Replace these defaults with your trace data when you have it. Year two and year three apply a 25% annual blended model-efficiency improvement to the model-driven portion of the cost curve.

Read the Build vs Buy guide
Architecture Calculator default Why the cost profile changes
ReAct / sequential prompt-chain agent 8 model steps, 6,500 tokens per step, 1.35 retry factor ReAct-style systems interleave reasoning, tool actions, observations, and follow-up reasoning. That loop tends to replay instructions, context, and observations across multiple model calls. ReAct paper · LangChain agent docs
RPA plus LLM workflow 6 model steps, 4,500 tokens per step, 1.25 retry factor Workflow systems usually have more fixed control flow than open-ended agents, but LLM steps still add extraction, classification, review, and exception loops. This is modeled below ReAct and multi-agent systems on model loops, but with an annual automation-platform software reserve because enterprise RPA paths commonly require platform, orchestrator, and bot-runner licensing. UiPath agents and workflows · Automation Anywhere pricing reference
Custom multi-agent framework 10 model steps, 8,500 tokens per step, 1.45 retry factor Multi-agent systems coordinate separate agent roles, tools, messages, and handoffs. That can improve capability, but it usually increases prompt volume, state carried forward, observability work, and validation paths. AutoGen framework docs · Dynamic reasoning cost study
Deterministic workflow plus LLM calls 6 model steps, 5,000 tokens per step, 1.25 retry factor Planning, DAG, and compiled execution patterns can reduce repeated model calls and context replay. The calculator still models this as a serious production build: teams must maintain eval harnesses, compiler/workflow versions, policy checks, observability, exception handling, and long-term governance. ReWOO paper · LLMCompiler paper
BLS: software engineering labor cost The engineering-cost input anchors against BLS software role wages, then adds senior mix, benefits, recruiting, equipment, and management load.
AWS and Google Cloud: raw workload infra is usage-based Lambda, Step Functions, Cloud Run, Pub/Sub, S3, and Cloud Observability are metered by requests, duration, transitions, storage, logs, metrics, or traces. At this default volume, the calculator treats raw cloud cost as modest and models token spend separately.
Production controls: security and observability add overhead The infrastructure line reserves budget for the production layer around the agent: logs, metrics, traces, posture management, threat monitoring, audit evidence, and operational support.
RPA platforms: software licensing is a separate cost layer When the RPA plus LLM architecture is selected, the model adds a separate annual platform-software reserve for orchestration, bot runners, and enterprise automation tooling rather than treating the path as cloud infrastructure plus tokens only.
Deloitte: token economics affect TCO Deloitte argues that AI economics increasingly need token-level TCO planning, FinOps discipline, and infrastructure strategy because usage can scale faster than unit prices fall. Read Deloitte analysis
Stanford and Artificial Analysis: inference cost keeps falling AI Index and independent benchmark data show rapid inference-price and capability improvement. The model uses a conservative 25% annual efficiency gain in years two and three rather than assuming flat model economics.
HPCA 2026: dynamic reasoning increases cost variance A system-level study of AI agents found that multi-step reasoning introduces material resource usage, latency variance, and infrastructure-cost tradeoffs. Read the study
AgentDiet: trajectories create token waste AgentDiet found that reducing redundant agent trajectory context can cut input tokens by 39.9%-59.7% and total computational cost by 21.1%-35.9% in evaluated coding-agent tasks. Read the paper
AIIA and ClearML: hidden GenAI costs are under-modeled Their enterprise survey found many teams underestimated total ownership cost, including usage growth, infrastructure, API instability, and operational burden. Read the survey report
LLM serving research: scale changes infrastructure needs KV-cache and serving research shows that context length, concurrency, throughput, and memory pressure change the infrastructure footprint as production volume grows. Read the KV-cache survey
tau-bench: real-world task success remains uneven tau-bench found that state-of-the-art function-calling agents can complete fewer than half of tasks in realistic tool and policy interaction domains. Read tau-bench
Below-baseline agents add human correction work When the selected build architecture is modeled below the 80% human-team baseline, the calculator adds correction and feedback labor that scales with workflow volume and model-backed steps.
NIST: security monitoring is recurring work NIST's information security continuous monitoring guidance supports treating security controls, evidence, and risk visibility as an ongoing operating requirement, not just a launch task. Read NIST SP 800-137

FAQ

AI Agent ROI & TCO Calculator FAQ

What is an AI agent TCO calculator?

An AI agent TCO calculator estimates the total cost of ownership for deploying and operating AI agents over time. A complete model should include human labor baseline, implementation labor, model and token usage, cloud and memory infrastructure, integration work, security controls, observability, evaluations, residual human review, maintenance, and platform costs.

How do you calculate AI agent ROI?

AI agent ROI is calculated by comparing the human-only operating baseline against the full automation TCO. A practical formula is: AI agent ROI = (human-only baseline cost - automation TCO) / automation TCO. The business-case version is: ROI = avoided labor, faster cycle time, reduced rework, and added throughput minus implementation, platform, model, infrastructure, governance, and residual review costs. For production agent systems, automation TCO should include agent accuracy, residual human review, engineering headcount, implementation timeline, model and token usage, cloud infrastructure, evaluation, observability, security, compliance, integrations, and maintenance.

What is the best way to measure AI agent ROI?

The best way to measure AI agent ROI is to start with a measured human baseline, then subtract the full production cost of automation. Use real workflow volume, average handling time, loaded labor cost, implementation cost, model and token usage, cloud infrastructure, integrations, compliance controls, evaluation, monitoring, residual human review, and ongoing maintenance. That keeps the ROI case tied to operating reality instead of a prototype demo.

How do I prove ROI on AI automation?

To prove ROI on AI automation, document the before-and-after operating model. Capture baseline volume, human minutes per workflow, reviewer cost, cycle time, rework rate, and exception rate. Then compare those numbers with the automated path, including launch delay, engineering effort, platform cost, token spend, infrastructure, accuracy, human review, governance, and support. A defensible ROI proof should show savings, payback period, residual risk, and the assumptions behind each number.

How should financial services teams calculate AI agent ROI?

Financial services teams should calculate AI agent ROI with compliance and control costs included. In lending, payments, insurance, banking, and capital markets workflows, the business case should include evidence handling, policy checks, audit trails, access control, human approval gates, exception routing, model evaluation, and legal or compliance review. A low-friction demo can look profitable while the production control layer changes the true TCO.

Is this AI agent ROI calculator free?

Yes. The browser calculator is free to use and does not require sign-up before you model assumptions. You can adjust the cost, volume, architecture, token, infrastructure, governance, and review inputs directly on the page.

How much does it cost to build an AI agent platform in-house?

The cost depends on team size, timeline, architecture, and workflow complexity. A regulated production build often requires 5 to 8 senior engineers for 12 to 18 months, plus ongoing maintenance, evaluation, monitoring, security, integration, and model-upgrade work.

How do I estimate AI agent implementation cost?

AI agent implementation cost should include engineering FTE, delivery timeline, data preparation, integrations, security review, evaluation tooling, observability, infrastructure, model usage, governance, maintenance, and residual human review. The calculator uses those inputs to estimate the cost to build AI agent infrastructure internally and compare it with a platform path.

Why are production AI agents more expensive than prototypes?

A prototype can be a prompt chain. Production needs policy control, document processing, source evidence, retries, exception routing, audit trails, access control, regression testing, observability, and model-upgrade governance. Those layers usually cost more than the initial demo.

How do token costs affect AI agent TCO?

Agent systems can call models many times for one workflow. Sequential architectures may replay context, tool definitions, documents, and validation steps repeatedly. That means token cost scales with workflow volume, retry rate, context size, and architecture choice. The calculator now applies a 25% annual efficiency improvement in years two and three to reflect lower model prices, fewer effective steps, fewer retries, and lower token load over time.

How does agent architecture affect AI agent ROI and TCO?

Agent architecture changes how many model calls, tool calls, retries, handoffs, and validation loops are needed for one completed workflow. ReAct and multi-agent systems tend to replay more context across steps, while deterministic workflows plus targeted LLM calls can reduce repeated model calls. Architecture also affects accuracy, residual review, infrastructure, observability, governance, and maintenance cost, so it directly changes both ROI and 3-year TCO.

Why do infrastructure costs change when volume increases?

Higher workflow volume increases storage, logging, traces, queueing, monitoring, evaluation runs, incident handling, and throughput requirements. At moderate volume, raw AWS or Google Cloud compute is usually not the main cost driver; the larger recurring cost is the production operating stack around the workload. Model token pricing is modeled separately.

What costs should AI agent total cost of ownership include?

AI agent total cost of ownership should include development labor, implementation timeline, data preparation, system integrations, identity and access controls, security review, audit logging, observability, compliance evidence, deployment paths, exception handling, model usage, memory and state stores, cloud infrastructure, evaluation tooling, maintenance, and residual human review.

Why do token assumptions differ by agent architecture?

Different agent architectures create different model-call patterns. ReAct and multi-agent systems often repeat context, tool outputs, and validation loops across more model calls. Deterministic workflow-plus-LLM systems can reduce repeated model calls, but a regulated internal build still needs eval harnesses, workflow versioning, policy checks, observability, exception handling, and governance. Year two and year three apply a 25% blended improvement factor to the model-driven portion of the cost curve.

Can I replace the calculator defaults with my own numbers?

Yes. The defaults are planning assumptions for an early business case. For procurement, use your own loaded engineering cost, workflow volume, model pricing, production traces, integration budget, and infrastructure assumptions.

Should I build or buy an AI agent platform?

Buying is usually stronger when the workflow is regulated, document-heavy, policy-bound, high volume, and not itself the company's core agent infrastructure IP. For a build vs buy AI agents decision, compare the internal build timeline, engineering ownership, infrastructure, governance, residual review, and 3-year TCO against the platform path. Building can make sense when the platform is core product IP and the organization is ready to own every production layer indefinitely.

What hidden costs make internal AI agent builds expensive?

The hidden costs are usually not the first prompt chain. They are evaluation harnesses, workflow versioning, policy governance, observability, exception handling, access control, audit evidence, integrations, model upgrades, incident response, memory and state management, and the human review needed when agent accuracy is below the operating baseline.

Should I build internally or hire a systems integrator?

A systems integrator can help with discovery, implementation, change management, and integrations, but the company still needs to own the production architecture, policy logic, evaluation process, audit trail, security model, and long-term maintenance plan. MightyBot gives teams two stronger paths: build with your own team on the MightyBot platform, or build with MightyBot FDE (forward-deployed engineering) on top of the platform. If you use a custom integrator, include their fees plus the ongoing internal ownership cost.

Can my internal team build on MightyBot instead of building everything from scratch?

Yes. The MightyBot path does not require outsourcing the whole initiative. Your internal team can use the MightyBot platform for policy execution, evidence handling, audit trails, orchestration, and production controls while keeping ownership of workflow design and operating model decisions.

How is the MightyBot path estimated?

The calculator uses a task-based planning model derived from monthly physical workflow volume and measured MightyBot billable task volume. Tasks per workflow follow the selected complexity band, from 2026 production billing observations: roughly 5 for simple operational workflows, 20 for standard document workflows, and 60 for complex multi-document quality gates, plus a one-time deployment setup estimate. It is a planning model, not a quote.

How is build-agent accuracy modeled?

The build-agent accuracy card is a planning assumption based on architecture choice, not a claim about any specific internal team. Open-ended ReAct and multi-agent systems are modeled lower because public agent benchmarks show task success varies widely under realistic tool use, retries, and policy constraints. Replace the default with your own eval results when available.

How are internal infrastructure costs estimated?

The calculator auto-estimates one-time setup and annual non-model infrastructure Opex from monthly workflow volume, model-backed workflow steps, internal engineering team size, and the selected agent architecture. The default annual estimate is intentionally modest, but it reserves budget for compute, orchestration, memory and state stores, observability, queues, storage, security tooling, evaluation tooling, monitoring, and operational support.

Is internal cloud and tooling an Opex cost?

Yes. Internal cloud, tooling, and AI infrastructure are modeled as recurring annual Opex. Internal data, security, and integration setup is modeled separately as a one-time implementation cost because that work happens before and during production launch.

Does prompt caching change AI agent TCO?

Yes, materially, and only on the input side. Cached input reads bill at roughly one tenth of the base input rate, so repeated-context agents save the most. Output tokens are never served from cache, which means heavily cached configurations become output-dominated. The calculator models this with a cache-attainment slider and separate input and output token prices. One caution: caches expire in minutes, so within-cycle caching is realistic while cross-run caching usually is not.

Can self-consistency voting make an agent as reliable as a deterministic workflow?

No. In a July 2026 measured study, single-pass agent evaluation flipped roughly 1 in 5 borderline decisions run to run. Majority voting across 3 independent passes cut that to roughly 10 percent while multiplying model spend by more than 3x, and it converges to an error floor rather than to zero. Voting also inherits shared blind spots: a check that silently fails to fire fails in every pass. Deterministic checks return the same verdict for the same input every time, which is why the calculator models voting as a cost multiplier with an accuracy floor rather than as a path to parity.

How many billable tasks does a workflow consume?

It depends on workload complexity. Production billing observations across deployed workflows in 2026 show roughly 5 billable tasks per workflow for simple operational cases, roughly 20 for standard document workflows, and roughly 60 for complex multi-document quality gates. The heaviest production compliance agents, such as covenant monitoring and SAR drafting, average 100 to 200 tasks per workflow. The calculator scales both the MightyBot task rate and the internal build’s token volume with the same complexity setting so the comparison stays symmetric. The task rate also scales with the model-backed steps per workflow you set, against an 8-step reference workflow, so a heavier or lighter workflow shape moves both sides of the comparison.

What did MightyBot measure in the July 2026 study?

MightyBot reconstructed a multi-document quality-gate workload on a fully synthetic corpus of 11 design documents and ran every configuration on the same model at temperature 0, counting tokens from the API usage records. Measured results per evaluation cycle: single-pass multi-viewpoint evaluation cost about $1.34 and hit output caps, a growing-context agentic configuration cost about $4.64 with a 3.6x input-replay multiplier, and a parity attempt with voting and verification cost $10.63 to $15.27 while still failing determinism and audit requirements. The full method and raw token tables are published in the agent evaluation cost study.

Why does the calculator split input and output token prices?

Because they behave differently. Input tokens dominate agent workloads, are roughly 5x cheaper than output tokens at list price, and are the only part prompt caching can reduce. Output tokens are typically around 9 percent of total tokens on document evaluation workloads but are never cached, so they dominate the cost of cached configurations and of any strategy that multiplies passes, such as self-consistency voting.

Is this an AI agent pricing calculator or a TCO calculator?

Both. It prices the MightyBot platform path with task-based pricing at your volume, and computes the full 3-year total cost of ownership of building the same capability internally: engineering, infrastructure, model tokens, quality strategy, and residual human review. Use it as an AI agent pricing calculator, a build-vs-buy TCO calculator, or a cost-per-task estimator.

How do I estimate AI agent cost per task?

Divide the annual platform or build cost by annual task volume. The calculator does this automatically: the platform side uses complexity-banded billable tasks per workflow against tiered per-task rates, and the build side converts token, engineering, and infrastructure spend into an equivalent per-case figure you can compare directly.

How does speed affect AI agent ROI?

Cycle time is a third ROI axis alongside labor and technology cost. When a regulated review clears in one day instead of three, the business banks two days per case: in lending that is two days of interest carry on the full loan value, in insurance it is SLA headroom instead of penalties, and in operations it is working capital released sooner. The calculator prices this as an optional speed dividend: enter your current turnaround in days and what one day faster is worth per case, and each path earns that dividend only on cases that clear without returning to the human review queue, from the month it goes live. MightyBot cases clear in minutes (a measured evaluation workload ran 50 to 105 seconds per cycle; complex multi-document cases run longer) with roughly 2% residual review, so nearly every case captures the dividend from go-live. A build path waits out its own build timeline first, then forfeits the dividend on every case its residual review rate sends back to the human queue.

Should an internal build close its accuracy gap with human review or with voting?

They are different cost curves with the same ceiling. Residual human review keeps model spend flat but converts every unresolved error into ongoing labor: at 58 percent accuracy, roughly two thirds of the human baseline persists as review work in production years. Self-consistency voting buys accuracy with compute: 3 passes lift a 58 percent architecture to about 78 percent for roughly 3.3x model spend, and because labor usually dwarfs tokens at production volume, voting is often net cheaper for weak architectures. The catch is at the top: voting floors near a 10.4 percent borderline error rate, so a strong 84 percent architecture only reaches about 89.6 percent while still paying 3.3x compute and 3.3x cycle time. Voting stops paying off exactly where architectures get good. Both paths flatten into a floor below production-grade accuracy, and neither reaches the 99 percent plus that governed deterministic execution delivers.

How fast do production agent workflows run per case?

On a measured multi-document evaluation workload, MightyBot runs completed in 50 to 105 seconds end to end, including checks, policy evaluation, and rationale generation. A comparable multi-pass agent configuration runs minutes to tens of minutes per case, because each voting pass repeats the full evaluation. Production cycle time scales with case complexity: heavier multi-document cases run into the tens of minutes on any architecture. Cycle time matters most for iterative pipelines where the gate runs on every fix cycle.