Live Last run · 2026-07-18 17:59 UTC · 37 models · 8 providers · 5 tasks

State of the FleetA living public benchmark of the AI models behind a multi-agent fleet — versioned, objectively graded, and documented in the open. The question isn't “which model is best?” It's “which model do I route each job to?”

This page is a live experiment, documented in public. Every result is fingerprinted to its task, grader, and harness version; nothing is hand-curated. When a grader gets stricter, old results are superseded and the audit trail records it.

fleet.arnao.ai Author · Byron Arnao Generated 2026-07-18 17:59 UTC
The stack, in motion

How the fleet reaches its models

There is no central router: each agent holds its own provider credentials and a tiered fallback chain, connecting directly to the model APIs it is allowed to use. The bubbles below are live request/response traffic along each agent's primary model path — this is the wiring the rest of the page benchmarks.

HARDWAREAGENTSPROVIDERS Mac Mini<host-a> · Gia+Mia+Tia+Cia containers · OllamaWindows PC<host-b> · Zia runtime Giafable 5 · full provider deckMiagemini-flash · opsTiagemma4 local · T2 capCiaqwen3-coder · local-onlyZiasonnet · WhatsApp · Win Anthropicfable · opus · sonnet · haikuGooglegemini 2.5 / 3.1OpenAIgpt-5.x · 4.1 · o-seriesxAIgrok-4.xOpenRouterrouted · kimi-k3 · glm · 400+Ollamalocal · gemma · qwen · $0 primary pathdirect fallbackrouted (OpenRouter)requestresponse

Reading top-to-bottom: Hardware (Mac Mini, Windows PC) hosts the agent containers, and every agent connects directly to its providers. Gia (Fable 5 primary) carries the full deck: direct Anthropic, Google, OpenAI, xAI and DeepSeek keys, plus OpenRouter as a routed path (how kimi-k3 and glm are reached) and local Ollama. Mia runs Gemini-flash with Anthropic and Ollama fallbacks; Tia is local-first and capped at T2 (Haiku / Gemini-flash); Cia is local-only by design — no cloud keys at all; Zia runs Sonnet on the Windows box. Solid edges are primary model paths, dashed grey edges are direct fallback tiers.

The fleet right now

Winners — current run

Auto-computed from the latest non-superseded results only. “Free” means a model that truly runs at $0 on the local node — grok variants are pricing-unconfirmed and are never labeled free.

Top model
gemini-3.1-flash-lite
Best overall — fastest-correct across 4 graded tasks (avg speed-rank 2.0).
Best free / local
llama3.1:8b
Truly free (runs on the local node). 2 correct across tested tasks, $0 marginal cost.
Best vision
grok-4.20-non-reasoning
Fastest correct image read: 0.91s (xAI).
Lowest cost
gpt-4.1-nano
Correct at $0.000013/call — the cloud price floor (code task).
Most efficient
gpt-4.1-nano
Best correct-per-dollar-second: $0.000013 in 0.86s (code).
Executive overview

This cycle ran 37 models across 8 providers through 5 task types under one harness. The headline: correctness has largely collapsed as a differentiator — math is solved fleet-wide and code is near-saturated — so the real signal now lives in speed, cost, and instruction-following discipline. The most discriminating task was the 50-word summary: only 21/36 models held the exact word-count budget — where the ability to obey a hard constraint, not raw intelligence, shows up.

Best overall this run was gemini-3.1-flash-lite; the cautionary tail was phi4:latest, which missed 3 graded tasks. My routing read, by objective: optimize for raw costgpt-4.1-nano at $0.000013/call; for speedgemini-3.1-flash-lite; for visiongrok-4.20-non-reasoning (0.91s); for local / $0llama3.1:8b; for best price-per-correct-secondgpt-4.1-nano. That is the actual decision I encode in the router: not a single ‘best’ model, but the right model dispatched per job-type.

I run this in public as a deliberate, controlled experiment — same prompts, versioned graders, results fingerprinted and superseded when a grader tightens. The ‘experiment’ framing is intentional: it lets me probe a fast-moving market through my own infrastructure without overfitting to any vendor’s benchmark theater. The goal isn’t to crown a winner — it’s a defensible, reproducible map of where each model earns its slot.

Responsible-AI lens: as correctness converges, governance — not capability — becomes the differentiator. Objective graders, a full audit trail, and per-job routing are what observable, well-governed agent deployment looks like: every answer is attributable to a known model, grader, and harness version, and a stricter test retroactively invalidates stale results instead of quietly inflating a score. That auditability is the prerequisite for trusting an autonomous fleet with real work.

37
models, validated live
8
providers
5
task types
567
total results logged
Per-task scorecard

How each task shook out

A compact read per task: how many models were tested, how many passed the objective grader, and the fastest & cheapest model that got it right.

Math37 models
37/37 correct
gemini-2.5-flash-lite · 0.65s
$ gpt-4.1-nano · $0.000027
37/37 correct — math is effectively solved across the fleet; speed & cost decide.
Code36 models
34/36 correct
grok-4.20-non-reasoning · 0.66s
$ gpt-4.1-nano · $0.000013
34/36 passed all 6 unit tests. The 2 misses are real logic/edge-case failures.
Summarize36 models
21/36 correct
gpt-4.1-nano · 0.76s
$ gpt-4.1-nano · $0.000016
Hard target: only 21/36 models hit the 50±2-word count. Most over- or under-shoot.
Creative37 models
qualitative
▲ —
$ — (no priced-correct model)
Qualitative task — 37 models produced a poem; correctness not objectively scored.
Vision18 models
17/18 correct
grok-4.20-non-reasoning · 0.91s
$ gemini-2.5-flash · $0.000028
17/18 read the image correctly. Only text-capable multimodal models were tested.
Visual read

Cost, correctness, latency

Lightweight, dependency-free charts rendered inline.

Math: cost vs latency

$1e-5$1e-4$1e-3$1e-2$1e-10s19s38sgemini-2.5-flash-lite — $0.000037, 0.65sgemini-3.1-flash-lite — $0.000039, 0.91sgpt-4o — $0.000965, 1.19sgpt-5.4-mini — $0.000154, 1.24sgrok-4.20-non-reasoning — $0.001911, 1.26sgpt-4.1-nano — $0.000027, 1.32sclaude-haiku-4-5 — $0.000613, 1.52sgpt-4.1-mini — $0.000038, 1.64sgpt-5.4 — $0.001428, 1.65sgpt-4o-mini — $0.000053, 1.74sgpt-4.1 — $0.001356, 1.76sclaude-opus-4-8 — $0.008520, 2.42so3-mini — $0.001238, 2.45sgrok-4.3 — $0.002280, 2.62so4-mini — $0.001339, 3.27sgemini-2.5-flash — $0.000031, 3.47sgpt-5.5 — $0.003265, 3.62sclaude-sonnet-4-6 — $0.002619, 3.64sgpt-5.6-sol — $0.004055, 3.64so3 — $0.009770, 4.80sclaude-fable-5 — $0.011730, 5.08sgrok-4.20-reasoning — $0.002412, 5.61sglm-5.2 — $0.002383, 5.83sinkling — $0.001852, 7.03sgemini-3.1-pro-preview — $0.001855, 8.18sgemini-2.5-pro — $0.000852, 8.21skimi-k3 — $0.004560, 11.83sqwen3.7-max — $0.007897, 34.37sgemini-2.5-flagemini-3.1-flagpt-4ogpt-5.4-minigrok-4.20-non-gpt-4.1-nanoclaude-haiku-4gpt-4.1-minigpt-5.4gpt-4o-minigpt-4.1claude-opus-4-o3-minigrok-4.3o4-minigemini-2.5-flagpt-5.5claude-sonnet-gpt-5.6-solo3claude-fable-5grok-4.20-reasglm-5.2inklinggemini-3.1-progemini-2.5-prokimi-k3qwen3.7-max cost per call (log) latency (s)
Correct cloud models on the math task. Down-and-left is cheaper and faster.

Correctness by provider

Thinking MachinesThinking Machines: 4/44/4 · 100%MoonshotMoonshot: 4/44/4 · 100%AlibabaAlibaba: 3/33/3 · 100%OpenAIOpenAI: 28/3028/30 · 93%GoogleGoogle: 15/1715/17 · 88%AnthropicAnthropic: 11/1311/13 · 85%xAIxAI: 10/1210/12 · 83%LocalLocal: 34/4434/44 · 77%
Share of gradable tasks answered correctly, aggregated per provider.

Code task: latency spread (cloud)

● median ○ p95 (cloud only)grok-4.20-non-reasoninggrok-4.20-non-reasoning median 0.66sp95 0.72s0.66sgemini-3.1-flash-litegemini-3.1-flash-lite median 0.71sp95 0.82s0.71sgpt-5.4-minigpt-5.4-mini median 0.75sp95 0.77s0.75sgpt-4.1-nanogpt-4.1-nano median 0.86sp95 1.37s0.86sgpt-4ogpt-4o median 0.97sp95 1.63s0.97sgpt-4.1-minigpt-4.1-mini median 1.08sp95 1.09s1.08sgpt-4.1gpt-4.1 median 1.14sp95 1.15s1.14sgpt-4o-minigpt-4o-mini median 1.15sp95 2.01s1.15sgpt-5.4gpt-5.4 median 1.25sp95 1.42s1.25sclaude-haiku-4-5claude-haiku-4-5 median 1.76sp95 2.04s1.76sclaude-sonnet-4-6claude-sonnet-4-6 median 1.98sp95 4.32s1.98sclaude-opus-4-8claude-opus-4-8 median 2.48sp95 5.54s2.48sgpt-5.6-solgpt-5.6-sol median 2.53sp95 2.92s2.53sgpt-5.5gpt-5.5 median 2.92sp95 3.03s2.92sgemini-2.5-flashgemini-2.5-flash median 3.48sp95 4.26s3.48sclaude-fable-5claude-fable-5 median 4.08sp95 4.24s4.08sgrok-4.20-reasoninggrok-4.20-reasoning median 4.60sp95 4.79s4.60sgrok-4.3grok-4.3 median 4.83sp95 10.30s4.83sinklinginkling median 5.18sp95 7.52s5.18sgemini-3.1-pro-previewgemini-3.1-pro-preview median 5.85sp95 5.97s5.85sgemini-2.5-progemini-2.5-pro median 10.38sp95 12.98s10.38skimi-k3kimi-k3 median 12.05sp95 37.43s12.05sqwen3.7-maxqwen3.7-max median 22.55sp95 24.85s22.55s
Median (filled) vs p95 (hollow) wall-clock per cloud model. Red = failed unit tests.
Full results

Every model, every task

Current (non-superseded) results, sorted by median time. Local models are muted; local $0 marks truly-free models, pricing TBD marks unconfirmed grok pricing. Hover a ✓/✗ for grader detail.

Math — 37 models, ranked by speed

ModelProviderMedian timeCost/callCorrect
gemini-2.5-flash-liteGoogle0.65s$0.000037
gemini-3.1-flash-liteGoogle0.91s$0.000039
gpt-4oOpenAI1.19s / p95 1.63s$0.000965
gpt-5.4-miniOpenAI1.24s / p95 2.54s$0.000154
grok-4.20-non-reasoningxAI1.26s / p95 1.41s$0.0019
gpt-4.1-nanoOpenAI1.32s / p95 1.81s$0.000027
claude-haiku-4-5Anthropic1.52s / p95 1.58s$0.000613
gpt-4.1-miniOpenAI1.64s / p95 2.87s$0.000038
gpt-5.4OpenAI1.65s / p95 1.71s$0.0014
gpt-4o-miniOpenAI1.74s / p95 1.79s$0.000053
gpt-4.1OpenAI1.76s / p95 2.20s$0.0014
claude-opus-4-8Anthropic2.42s / p95 2.54s$0.0085
o3-miniLocal2.45s / p95 2.55s$0.0012
grok-4.3xAI2.62s / p95 3.56s$0.0023
o4-miniLocal3.27s / p95 3.67s$0.0013
gemini-2.5-flashGoogle3.47s / p95 4.00s$0.000031
gpt-5.5OpenAI3.62s / p95 4.18s$0.0033
claude-sonnet-4-6Anthropic3.64s / p95 4.14s$0.0026
gpt-5.6-solOpenAI3.64s / p95 3.84s$0.0041
o3Local4.80s / p95 5.84s$0.0098
claude-fable-5Anthropic5.08s / p95 5.33s$0.0117
grok-4.20-reasoningxAI5.61s / p95 5.77s$0.0024
glm-5.2Local5.83s / p95 11.05s$0.0024
inklingThinking Machines7.03s / p95 7.81s$0.0019
llama3.1:8bLocal7.07s$0
gemma3:27bLocal7.85s$0
gemini-3.1-pro-previewGoogle8.18s / p95 8.91s$0.0019
gemini-2.5-proGoogle8.21s / p95 8.93s$0.000852
kimi-k3Moonshot11.83s / p95 21.38s$0.0046
gemma3:12bLocal12.62s$0
gemma4:latestLocal13.26s$0
qwen3-coder:30bLocal16.23s$0
qwen2.5-coder:32bLocal21.38s$0
phi4:latestLocal22.55s$0
deepseek-r1:14bLocal23.75s$0
mistral-small3.2Local24.62s$0
qwen3.7-maxAlibaba34.37s / p95 37.87s$0.0079

Code — 36 models, ranked by speed

ModelProviderMedian timeCost/callCorrect
grok-4.20-non-reasoningxAI0.66s / p95 0.72spricing TBD
gemini-3.1-flash-liteGoogle0.71s / p95 0.82s$0.000031
gpt-5.4-miniOpenAI0.75s / p95 0.77s$0.000105
gpt-4.1-nanoOpenAI0.86s / p95 1.37s$0.000013
gpt-4oOpenAI0.97s / p95 1.63s$0.00075
gpt-4.1-miniOpenAI1.08s / p95 1.09s$0.000026
gpt-4.1OpenAI1.14s / p95 1.15s$0.00056
gpt-4o-miniOpenAI1.15s / p95 2.01s$0.000043
gpt-5.4OpenAI1.25s / p95 1.42s$0.0012
claude-haiku-4-5Anthropic1.76s / p95 2.04s$0.0014
claude-sonnet-4-6Anthropic1.98s / p95 4.32s$0.0028
o3-miniLocal2.09s / p95 2.14s$0.0015
claude-opus-4-8Anthropic2.48s / p95 5.54s$0.0073
gpt-5.6-solOpenAI2.53s / p95 2.92s$0.0018
o3Local2.67s / p95 4.17s$0.0089
gpt-5.5OpenAI2.92s / p95 3.03s$0.0021
o4-miniLocal3.27s / p95 3.33s$0.0018
glm-5.2Local3.33s / p95 4.91s$0.0016
gemini-2.5-flashGoogle3.48s / p95 4.26s$0.000161
claude-fable-5Anthropic4.08s / p95 4.24s$0.005
grok-4.20-reasoningxAI4.60s / p95 4.79spricing TBD
grok-4.3xAI4.83s / p95 10.30spricing TBD
inklingThinking Machines5.18s / p95 7.52s$0.001
gemini-3.1-pro-previewGoogle5.85s / p95 5.97s$0.0012
phi4:latestLocal9.48slocal $0
gemini-2.5-proGoogle10.38s / p95 12.98s$0.0013
gemma4:latestLocal12.04slocal $0
kimi-k3Moonshot12.05s / p95 37.43s$0.0033
gemma3:12bLocal13.90slocal $0
qwen3-coder:30bLocal15.42slocal $0
qwen2.5-coder:32bLocal18.05slocal $0
llama3.1:8bLocal19.01slocal $0
mistral-small3.2Local20.10slocal $0
qwen3.7-maxAlibaba22.55s / p95 24.85s$0.0037
gemma3:27bLocal25.08slocal $0
deepseek-r1:14bLocal66.31slocal $0

Summarize — 36 models, ranked by speed

ModelProviderMedian timeCost/callCorrect
grok-4.20-non-reasoningxAI0.62s / p95 0.67spricing TBD
gpt-4.1-nanoOpenAI0.76s / p95 0.88s$0.000016
gemini-3.1-flash-liteGoogle0.79s / p95 0.81s$0.000031
gpt-5.4-miniOpenAI1.00s / p95 1.09s$0.000144
gpt-4.1OpenAI1.00s / p95 1.40s$0.000634
claude-haiku-4-5Anthropic1.04s / p95 2.85s$0.000396
gpt-4oOpenAI1.24s / p95 1.40s$0.000733
gpt-4o-miniOpenAI1.26s / p95 1.40s$0.000046
gpt-5.4OpenAI1.46s / p95 1.47s$0.0013
gpt-4.1-miniOpenAI1.55s / p95 2.65s$0.000031
claude-opus-4-8Anthropic2.45s / p95 2.61s$0.0108
o3-miniLocal2.66s / p95 4.28s$0.0038
claude-sonnet-4-6Anthropic2.81s / p95 3.08s$0.0015
o3Local3.26s / p95 3.32s$0.022
gemini-2.5-flashGoogle4.15s / p95 7.20s$0.000022
grok-4.20-reasoningxAI4.70s / p95 15.68spricing TBD
gpt-5.5OpenAI4.77s / p95 5.45s$0.0065
o4-miniLocal4.80s / p95 8.80s$0.0024
gpt-5.6-solOpenAI5.08s / p95 5.88s$0.0089
llama3.1:8bLocal6.20slocal $0
claude-fable-5Anthropic7.58s / p95 8.99s$0.0297
grok-4.3xAI8.98s / p95 9.03spricing TBD
phi4:latestLocal9.51slocal $0
gemma3:12bLocal10.52slocal $0
inklingThinking Machines14.93s / p95 104.58s$0.0245
qwen3-coder:30bLocal15.53slocal $0
gemini-2.5-proGoogle15.66s / p95 16.77s$0.000663
deepseek-r1:14bLocal17.40slocal $0
glm-5.2Local18.32s / p95 22.97s$0.0061
gemma4:latestLocal18.62slocal $0
gemma3:27bLocal19.89slocal $0
qwen2.5-coder:32bLocal19.94slocal $0
mistral-small3.2Local21.11slocal $0
gemini-3.1-pro-previewGoogle28.49s / p95 33.88s$0.0011
kimi-k3Moonshot35.45s / p95 64.25s$0.0172
qwen3.7-maxAlibaba56.06s / p95 79.23s$0.0136

Creative — 37 models, ranked by speed

ModelProviderMedian timeCost/callCorrect
gemini-2.5-flash-liteGoogle0.55s$0.000018
gpt-4.1-nanoOpenAI0.67s / p95 0.82s$0.000009
gemini-3.1-flash-liteGoogle0.87s / p95 0.94s$0.00002
grok-4.20-non-reasoningxAI0.87s / p95 0.94s$0.001
gpt-4o-miniOpenAI1.00s / p95 1.39s$0.000031
gpt-4.1-miniOpenAI1.00s / p95 1.14s$0.00002
gpt-5.4-miniOpenAI1.08s / p95 1.34s$0.000094
claude-haiku-4-5Anthropic1.23s / p95 1.28s$0.000279
gpt-4.1OpenAI1.26s / p95 1.53s$0.000294
gpt-4oOpenAI1.34s / p95 1.38s$0.000508
gpt-5.4OpenAI2.20s / p95 2.34s$0.00092
claude-sonnet-4-6Anthropic2.56s / p95 2.60s$0.0011
claude-opus-4-8Anthropic3.34s / p95 4.35s$0.0063
o3Local4.97s / p95 9.65s$0.0206
grok-4.3xAI5.40s / p95 7.30s$0.000972
llama3.1:8bLocal5.73s$0
o3-miniLocal7.09s / p95 7.69s$0.0066
gemini-2.5-flashGoogle7.20s / p95 8.11s$0.000014
phi4:latestLocal8.21s$0
gemma3:12bLocal10.00s$0
o4-miniLocal10.08s / p95 11.15s$0.0066
claude-fable-5Anthropic10.49s / p95 17.10s$0.0255
gpt-5.5OpenAI11.24s / p95 11.39s$0.0115
qwen3-coder:30bLocal14.80s$0
qwen2.5-coder:32bLocal15.69s$0
gemini-2.5-proGoogle15.74s / p95 22.97s$0.000381
grok-4.20-reasoningxAI16.01s / p95 17.39s$0.000909
gpt-5.6-solOpenAI16.33s / p95 26.15s$0.0171
mistral-small3.2Local17.96s$0
gemma3:27bLocal18.10s$0
gemma4:latestLocal24.63s$0
deepseek-r1:14bLocal29.19s$0
glm-5.2Local35.79s / p95 52.97s$0.006
inklingThinking Machines68.44s / p95 88.28s$0.0222
gemini-3.1-pro-previewGoogle95.62s / p95 104.65s$0.000633
qwen3.7-maxAlibaba148.44s / p95 161.79s$0.0176
kimi-k3Moonshot191.79s / p95 274.26s$0.065

Vision — 18 models, ranked by speed

ModelProviderMedian timeCost/callCorrect
grok-4.20-non-reasoningxAI0.91s / p95 0.93s$0.0038
gemini-3.1-flash-liteGoogle1.36s / p95 1.41s$0.000118
gpt-4oOpenAI1.72s / p95 1.82s$0.003
gpt-5.6-solOpenAI2.02s / p95 2.60s$0.0072
grok-4.3xAI2.15s / p95 2.28s$0.0042
gemini-2.5-flashGoogle2.29s$0.000028
gpt-5.5OpenAI2.45s / p95 16.35s$0.0068
grok-4.20-reasoningxAI2.62s / p95 2.72s$0.0038
inklingThinking Machines2.80s / p95 3.43s$0.0019
claude-fable-5Anthropic3.62s / p95 3.66s$0.0163
gemini-2.5-proGoogle5.46s / p95 5.82s$0.000538
gemini-3.1-pro-previewGoogle6.75s / p95 8.56s$0.003
phi4:latestLocal8.28s$0
gemma3:12bLocal10.54s$0
gemma3:27bLocal20.08s$0
kimi-k3Moonshot24.89s / p95 27.09s$0.0078
gemma4:latestLocal27.24s$0
mistral-small3.2Local29.85s$0
Methodology

The protocol

Every model sees the same prompt per task. Cloud models are timed over N=3 trials (median reported, p95 where available); local models are single-trial and their latency includes cold model-load. Grading is objective wherever possible.

Graders, in plain English

Hypotheses under test

Limitations — stated openly
  • Grok pricing unconfirmed. xAI variants are excluded from cost rankings and never labeled “free.” Speed is real; economics are not yet validated.
  • Local latency includes cold-load. “$0” ignores wall-clock (often 20–110s) and local compute.
  • N=3 cloud trials is enough for a stable median, not a tight tail estimate.
  • Creative is not objectively graded — it's qualitative and excluded from correctness math.
  • Vision tested only multimodal-capable models; text-only models were skipped by design, not failed.
Versioning & audit trail

Nothing is hand-curated

Each result carries a run_key = fingerprint of model + task + task_version + grader_version + harness_version. A newer, higher-version result for the same model+task supersedes the old one. 142 of 567 logged results are currently superseded — preserved for audit, excluded from the dashboard.

What changed — harness

What changed — graders

Invalidation triggers
  • Vision grader v1→v2 (accept word-numbers like “three”): superseded 28 vision results that were mis-scored on digit-only matching.
  • Code & Summarize grader v2 + harness v3 (objective execution / exact word-count, full-output capture): superseded 59 single-shot/eyeball results across code and summarize.

Run log

DateRecordsTasks run
2026-07-1821Code, Creative, Math, Summarize, Vision (2 since superseded)
2026-07-0992Code, Creative, Math, Summarize
2026-07-01107Code, Creative, Math, Summarize, Vision
2026-06-21105Code, Creative, Math, Summarize, Vision
2026-06-1456Creative, Math, Vision
2026-06-13158Code, Creative, Math, Summarize, Vision (114 since superseded)
2026-06-1128Code, Creative, Math, Summarize (26 since superseded)
Security posture

Documented in public, redacted by policy

This is a live experiment run on real infrastructure and shared openly. Infrastructure details (IP addresses, hostnames) are redacted with functional tags<host-a>, <local-node> — so the methodology is fully reproducible without exposing the network. Benchmark data, prompts, graders, and version history are public; the wiring is not.

Archive

The earlier write-ups

The original June 11 fleet analysis (architecture, tokenomics, 7-model v1 benchmark) and the June 13 v2 evolution narrative are preserved for continuity.

Open the v1 / v2 narrative archive

The v1 page asked “can the models do the math?” (7 cloud models, single prompt) and documented the fleet architecture and tokenomics. The v2 evolution expanded to 37+ models across 8 providers and added a vision task. Both have been superseded by this living, versioned page — which recomputes every winner directly from the append-only result store rather than from hand-written prose. The full prior narrative remains in version control.

Key v2 findings that still hold: math is solved across the fleet; the cloud price floor is effectively zero (sub-$0.0001 correct answers); local models are correct but slow (cold-load latency dominates); verbosity, not sticker price, drives cost.