~/coding-agent

Page generated data from

Coding Agent

Which coding-agent subscription, model and reasoning effort to use — and how long your quota lasts.

Data checked on September 28, 2026 · 117 sources

What to use, when

Pick what you are about to do: you get the provider, model, reasoning effort and subscription to use, with the evidence behind the choice.

What do you want to do?

Sorted from the most to the least frequent use of coding agents in 2026, according to OpenRouter, Anthropic Economic Index, Claude Code, AIDev.

Everyday coding and debugging

Features, debugging, bug fixes, reviews, small refactors — the work you do all day.

  • Fix a failing test or a reported bug
  • Add an endpoint, a form, a CLI option
  • Review a pull request
  • Small refactor in a few files
✓ Recommended

Claude Opus 5.5(default)·effort medium(default)

Agent
Claude Code
Subscription
Claude Pro ($20) or Max ($100 / $200)
Provider
Anthropic
Alternative

GPT-6 Sol·effort high

Codex · ChatGPT Plus ($20) or Pro ($100 / $200)

Codex on ChatGPT Plus: DeepSWE 65.3% at $0.64 per task — +8.6 points over medium for +$0.26; Plus allows about 15–150 Sol messages per 5 hours.

Alternative

GLM-5.3·effort max(default)

Claude Code or ZCode · GLM Coding Plan ($18 / $80 / $168)

Budget: GLM Coding Plan Lite ($18/month, 2,000 credits per 5 hours) inside Claude Code; DeepSWE 69.0% at max.

How to work

  1. No plan needed when you can describe the change in one sentence; otherwise type /plan <task> — more reliable than Shift+Tab, which now starts from auto mode.
  2. Stay on medium (the default). For a stubborn bug, raise it with /effort high, or add ultrathink to a single prompt.
  3. Give the agent a check it can run (tests, build, lint) and ask for the command and its output, not a claim. After two failed corrections, /rewind or /clear and re-prompt with what you learned.
  4. Before the pull request: /code-review looks for bugs in a fresh subagent; /simplify cleans up the diff but does not look for bugs.
  5. In Codex: /model → GPT-6 Sol, effort high; structure the prompt as Goal, Context, Constraints, Done when.

Why

  • Best quality per unit of quota: 51.2 on the AA Intelligence Index for $1.34 per task, ahead of GPT-6 Astra at high (50.9 for $1.73).
  • Anthropic's own SWE-bench Pro subset: 92.8% solved at $0.22 per solved task, versus 92.3% at $1.19 for Fable 5.1 at its default.
  • It is Claude Code's default model and effort since 2026-09-22. Step up to high for a hard bug: +4 points on Terminal-Bench 4.0 (52.5% → 56.6%) for about 27% more quota.

Sources: Artificial Analysis model leaderboard · Optimizing for cost and intelligence · Model configuration · DeepSWE v1.1 leaderboard · Introducing GPT-6 Sol and Luna · Codex pricing · GLM Coding Plan · Claude Code best practices · Claude Code commands · Claude Code permission modes

Habits that pay off in every agent

Whatever the task: what Claude Code, Codex, Antigravity and Grok Build document, and what Claude Code’s creator does, checked against current docs (outdated tips left out).

  1. 01

    Verification first

    Give the agent a pass/fail check (tests, build, screenshot) and ask for evidence, not claims. Every vendor puts this tip first.

    Sources: Claude Code best practices · Codex best practices · Antigravity CLI best practices · How I use Claude Code (thread)

  2. 02

    Plan only when it helps

    Plan ambiguous, multi-file or unfamiliar work; skip it for one-sentence changes. Recent models plan within the conversation, and /plan stays available.

    Sources: Claude Code best practices · Grok Build plan mode · Boris Cherny on plan mode (Hacker News)

  3. 03

    One task per session

    /clear between unrelated tasks, /rewind instead of piling up corrections, /btw for a side question, /compact <what to keep> mid-task.

    Sources: Claude Code best practices · Claude Code commands

  4. 04

    Short instruction files

    Keep AGENTS.md / CLAUDE.md short: add a rule when a mistake repeats, remove what the model already does, and re-test after each new model (/doctor).

    Sources: Claude Code best practices · Codex best practices

  5. 05

    Raise the effort, don't beg

    Start at the agent's default effort; raise it (or ultrathink for one turn) instead of writing “think harder”, and keep xhigh and max for tasks where they measurably help.

    Sources: Prompting Claude Opus 5.5 · Model configuration

  6. 06

    Fresh-context review

    Have code checked by an agent that did not write it (/code-review, /review), and ask it to report only correctness problems.

    Sources: Claude Code best practices · Claude Code commands

  7. 07

    Parallel work costs quota

    Worktrees, parallel sessions, subagents and workflows multiply the work done, but all draw on the same subscription limits: check the simulator before scaling out.

    Sources: Claude Code costs · Claude Code worktrees

  8. 08

    Auto mode, not skip-permissions

    Use auto mode (a classifier screens each action) or a sandbox; never --dangerously-skip-permissions in a normal session.

    Sources: Claude Code permission modes

What not to do right now

The habits that waste quota or quality with today’s models and plans — each with what to do instead and the figures behind it. Re-checked at every data update.

Models

  • Don’t: Code everything with Claude Fable 5.1, especially at max

    Instead: Claude Opus 5.5 at medium (default), xhigh for long runs.

    Opus 5.5 beats it everywhere for a fraction of the quota. Terminal-Bench 4.0 (AA): Fable max 52.0% for $19.22 per task, below its own xhigh (55.1%) and below Opus 5.5 xhigh (59.6% for $8.78). CursorBench 4.0: Fable max 51.8 for $17.28, Opus 5.5 high 56.0 for $3.97. On Max plans Fable is also capped at 50% of the weekly limit.

    Sources: Artificial Analysis model leaderboard · CursorBench 4.0 · Claude Fable models on your plan

  • Don’t: Pick Claude Sonnet 5 to save quota

    Instead: Claude Opus 5.5 at low for cheap work, or medium (default).

    Opus 5.5 at low beats Sonnet 5 at max on the AA index (42.3 vs 38.2) for a ninth of the cost ($0.55 vs $5.09), and Opus 5.5 is Claude Code's default model.

    Sources: Artificial Analysis model leaderboard · Model configuration

  • Don’t: Use GPT-6 Sol, even at max, for long autonomous terminal runs

    Instead: GPT-6 Astra at high: same cost, ten points higher.

    Terminal-Bench 4.0 (AA): Sol max 43.9% for $4.00 per task, Astra high 54.0% for $4.05. Sol stays a good everyday model.

    Sources: Artificial Analysis model leaderboard

  • Don’t: Give Kimi K3 long terminal-heavy tasks

    Instead: GLM-5.3 at max (41.9%) on a budget; keep Kimi K3 for web UI and data visualization, where it ranks near the top.

    Terminal-Bench 4.0 (AA): 12.6% at max, the same as at low.

    Sources: Artificial Analysis model leaderboard · Design Arena: data visualization

Reasoning effort

  • Don’t: Leave max on for everyday work

    Instead: The agent's default effort; xhigh for multi-hour runs; max only where you have measured a gain.

    Opus 5.5: medium (default) 54.6% on FrontierCode for $0.80, max 54.4% for $6.19; on Terminal-Bench 4.0 max equals xhigh (59.6%) for 49% more. GPT-6 Astra: xhigh matches or beats max on Terminal-Bench 4.0 and DeepSWE (59.6% vs 59.1%; 74.1% vs 73.2%) for 30–40% less. Anthropic warns that max can overthink.

    Sources: FrontierCode leaderboard data · Artificial Analysis model leaderboard · DeepSWE v1.1 leaderboard · Model configuration

  • Don’t: Drop Claude Opus 5.5 to low for real agentic work

    Instead: medium (default); keep low for subagents and bulk edits.

    Terminal-Bench 4.0 (AA): low 31.3% vs medium 52.5% — 21 points lost to save $1.97 per task.

    Sources: Artificial Analysis model leaderboard

  • Don’t: Push Grok 4.7 above high (default)

    Instead: Stay on high (default).

    xhigh adds 1 point on Terminal-Bench 4.0 (24.7% → 25.8%) for 63% more cost, and scores lower on FrontierCode (47.6% → 45.6%).

    Sources: Artificial Analysis model leaderboard · FrontierCode leaderboard data

Habits

  • Don’t: Run ultracode or a workflow for ordinary tasks

    Instead: One session for normal work; keep workflows for a large migration or audit, tried on one directory first.

    A workflow orchestrates dozens of agents, and every agent draws on the same subscription limits: Claude Code warns beyond 25 agents or 1.5M tokens. Combined with Fable or max, a single run can burn a week of quota.

    Sources: Claude Code workflows · Claude Code costs

  • Don’t: Run with --dangerously-skip-permissions to go faster

    Instead: Auto mode (the default): a classifier screens each action; or a sandbox or container.

    It removes every safeguard on file deletion, shell commands and network access, while auto mode already avoids most prompts.

    Sources: Claude Code permission modes · Claude Code best practices

  • Don’t: Force plan mode on every small change

    Instead: Plan only ambiguous or multi-file work with /plan.

    The docs say to skip the plan when the change fits in one sentence, and Claude Code's creator says plan mode is no longer useful with Opus 5.5 — it mostly costs a round trip.

    Sources: Claude Code best practices · Boris Cherny on plan mode (Hacker News)

How long will your subscriptions last?

Enter the subscriptions you have and how you work: you get the agent hours per 5-hour window and per week, and when the quota runs out.

Performance against quota burn

Each line is one model across its reasoning-effort levels. Further left = the task burns less of your quota; higher = better result. The thin step line is the efficient frontier: no other model × effort does better for less.

Benchmark

IndependentAgentic terminal tasks run by Artificial Analysis with the same harness for every model, at every effort level.

Source updated continuously · Retrieved Sep 28, 2026

Highlight models (thicker line and label; use the chart legend to hide models)

Score vs cost per task

API-equivalent cost at list prices, log scale — the same ratio at which subscription quotas burn.

Scroll or pinch to zoom (Shift + scroll for the vertical axis), drag the slider, or use the zoom tool at the top right. Click a legend entry to hide or show it.

17 results without a published cost are shown only in the effort chart and the table.

Score by reasoning effort

Where the line flattens, the next effort level costs more quota for little or no gain.

Show the data table

Coding leaderboards

Independent rankings of coding agents and web development. Switch each leaderboard on or off; scroll or zoom inside a chart, and click a provider in its legend to hide it.

Leaderboards to show

Artificial Analysis Coding Agent Index v1.5

Each coding agent with its vendor's model (Claude Code, Codex, Grok Build, Kimi Code, Antigravity, OpenCode) on DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA — the closest proxy for what a subscription gives you.

Source updated Sep 6, 2026 · Retrieved Sep 28, 2026 · 20 models

Incomplete coverage — this data does not include Claude Sonnet 5. Key current models are missing, so the ranking here is not fully relevant.

Show the data table

Arena Code: WebDev

Blind pairwise human votes on web-development tasks (795,514 votes, cutoff 2026-09-25).

Source updated Sep 25, 2026 · Retrieved Sep 28, 2026 · 134 models

Show the data table

Design Arena: agentic web app, full-stack

Two agents build the same full-stack app (React, Supabase, deployed to Vercel) in a sandbox with 23 tools; people vote for the better one.

Source updated Sep 28, 2026 · Retrieved Sep 28, 2026 · 49 models

Incomplete coverage — this data does not include Claude Opus 5.5, Grok 4.7, GLM-5.3. Key current models are missing, so the ranking here is not fully relevant.

Show the data table

Official launch figures

What each vendor published for its own model at launch, with the effort it used. Vendors pick their own agent, harness and settings, so these numbers usually run higher than independent measurements — use them to see each vendor's claims, and the charts above to compare.

ModelSWE-bench VerifiedSWE-bench ProTerminal-Bench 2.xTerminal-Bench 4.0DeepSWE v1.1FrontierCode 1.1
Claude Fable 5.1Anthropic—81.2%↗max—55.8%↗max—50.3%↗max
Claude Opus 5.5Anthropic—89.9%↗max—66.4%↗xhigh74.2%↗max54.4%↗max
Claude Sonnet 5Anthropic85.2%↗max63.2%↗max80.4%↗xhigh · 2.1———
Claude Haiku 4.5Anthropic73.3%↗—————
GPT-6 AstraOpenAI———57.88%↗high74.12%↗xhigh53.3%↗max
GPT-6 SolOpenAI————68.81%↗max49.27%↗max
GPT-6 LunaOpenAI————66.59%↗max42.42%↗max
Gemini 3.8 FlashGoogle——89.4%↗medium · 2.119.1%↗73.7%↗high (default)—
Gemini 3.7 FlashGoogle——85.8%↗medium (default) · 2.111.2%↗65.3%↗—
Gemini 3.1 Pro (preview)Google80.6%↗high (default)54.2%↗high (default)68.5%↗high (default) · 2.0———
Grok 4.7xAI———38%↗xhigh71%↗high (default)—
Grok 4.6xAI———20.3%↗high (default)65.9%↗high (default)—
Kimi K3Moonshot AI——88.3%↗max · 2.1—67.5%↗max—
GLM-5.3Z.ai——88.2%↗max (default) · 2.1—66.9%↗max (default)—
GLM-5.3-FlashZ.ai——84.3%↗2.1—63.4%↗—

Remote or local?

Beyond the six providers above: the other subscriptions that are worth it, and when running a model on your own machine makes sense.

Other subscriptions worth knowing

SubscriptionPrice / monthQuotaVerdict
GitHub Copilot↗↗$10 / $39 / $100Dollar credits at API prices: $15 on Pro, $70 on Pro+, $200 on Max; completions free.worth itSecond most used coding agent (21%); GPT-6, Opus 5.5, Fable 5.1, Gemini 3.8 Flash, Grok 4.7 and Kimi K3 in VS Code, JetBrains and the CLI.
Alibaba Qwen Coding Plan↗↗↗$506,000 requests per 5 hours, 45,000 per week, 90,000 per month.worth itPublished quotas; Qwen3.8 Max is #6 on Arena WebDev and 43.3 on the AA Coding Agent Index; works in Claude Code, Codex and OpenCode.
OpenCode Go↗$10Per 5 hours: Kimi K3 110 requests, Qwen3.8 Max 160, GLM-5.3 220, DeepSeek V4.1 Flash 26,000.worth itThe cheapest way to use the best open-weight models in a good agent (OpenCode is the harness AA uses for GLM).
MiniMax Token Plan↗↗↗$22 / $55 / $1325-hour and weekly windows, in tokens.watchClear plan and 47.2% on SWE-rebench, but only 2.0% on AA's Terminal-Bench 4.0 run: watch the next model.
Meta Muse Code↗↗$5 / $15 / $5010–50 prompts per 5 hours on Everyday; 5× and 20× above.niche54.3 on the AA Coding Agent Index (above Kimi K3) for a small price.
Xiaomi MiMo Token Plan↗↗$6 – $100Credits, 20% off during off-peak hours.nicheVery cheap; MiMo-V2.6-Pro reaches 34.8% on AA Terminal-Bench 4.0.
Cursor↗$20 / $60 / $200Two pools (Cursor models, other models), sizes not published.nicheAn IDE surface for many models, including Grok 4.7 and Composer 2.5.
Devin↗↗$20 / $200Not published.nicheDevin Fusion (Fable 5.1 + SWE-2) scores 61.7 on the AA Coding Agent Index.
Factory Droid↗$20 / $100 / $200Per-model multipliers (Opus 5.5 1.6×, GPT-6 Astra 4×).nicheMulti-model agent with explicit multipliers.
Mistral Vibe↗↗$14.99 / $24.99Not published.nicheEuropean option; Mistral Medium 3.5 scores 0% on AA's Terminal-Bench 4.0 run.

Running models locally

Local models are a separate tier: they cost no quota, but the best of them reach 5–33% on Terminal-Bench 4.0 versus about 60% for frontier models. Pair a local model for routine work with a remote frontier model for hard tasks.

HardwareRecommended modelMemoryDecode speedQuality (AA index · Terminal-Bench 4.0)
24 GB GPU (RTX 4090 / 3090)Qwen3.8-27B Q4_K_M16.5 GB38–46 tok/sAA 33.7 · TB4 5.6%↗↗↗
32 GB GPU (RTX 5090)Qwen3.8-27B Q6 / NVFP422–25 GB23–75 tok/s · MTP 98–134AA 33.7 · TB4 5.6%↗↗
128 GB unified memory (Strix Halo, DGX Spark, M5 Max)Qwen3.8-Flash-Next IQ4_XS93.7 GB11–37 tok/s · MTP 47–84AA 39.8 · TB4 25.3%↗↗
256 GB unified memory (M5 Ultra, M3 Ultra)GLM-5.3-Flash MLX 4-bit204 GB26–41 tok/sAA 41.8 · TB4 32.8%↗↗

Worth it for

  • Private or regulated code that must not leave the machine
  • Working offline
  • Hardware you already own
  • High-volume routine work: boilerplate, tests, subagents
  • A 256 GB machine shared by a team

Not worth it for

  • Saving money versus a $20 subscription (hardware and power cost $10–270 a month)
  • Hard agentic tasks: the quality gap to frontier models is large

References

Every figure on this page comes from these pages. For each one: what we took from it, and when.

  1. Model Studio Coding PlanAlibaba CloudOfficial

    Qwen Coding Plan Pro price and request quotas per 5 hours, week and month.

    Published Sep 11, 2026 · Accessed Sep 28, 2026

  2. Anthropic Economic Index: impact on software developmentAnthropicOfficial

    Languages (JS/TS 31%, HTML/CSS 28%) and task clusters (UI/UX components 12%, first) in 500,000 coding interactions.

    Published Apr 28, 2025 · Accessed Sep 28, 2026

  3. Claude Code /goalAnthropicOfficial

    `/goal <condition>`: a separate evaluator re-checks the condition after each turn.

    Accessed Sep 28, 2026

  4. Claude Code best practicesAnthropicOfficial

    Verification first (prompt, `/goal`, Stop hook, subagent); skip the plan for one-sentence changes; prune CLAUDE.md; `/clear` after two failed corrections; `/batch` fan-out; adversarial review.

    Accessed Sep 28, 2026

  5. Claude Code commandsAnthropicOfficial

    `/code-review` (alias `/review`), `/simplify` (cleanup, not bug finding), `/batch`, `/goal`, `/effort`, `/doctor`, `/btw`, `/rewind`.

    Accessed Sep 28, 2026

  6. Claude Code costsAnthropicOfficial

    Subagents and agent teams draw on the same usage limits; agent teams use about 7× more tokens when teammates run in plan mode.

    Accessed Sep 28, 2026

  7. Claude Code in ChromeAnthropicOfficial

    `claude --chrome` lets the agent drive the browser; requires a subscription login (`/login`), unavailable with an API key.

    Accessed Sep 28, 2026

  8. Claude Code permission modesAnthropicOfficial

    Auto mode is the starting mode since v2.1.283; `Shift+Tab` cycles auto → default → acceptEdits → plan; `/plan`.

    Accessed Sep 28, 2026

  9. Claude Code workflowsAnthropicOfficial

    `ultracode` keyword or “use a workflow”; quota cost and the warning beyond 25 agents or 1.5M tokens.

    Accessed Sep 28, 2026

  10. Claude Code worktreesAnthropicOfficial

    `claude --worktree <name>` / `-w` for isolated parallel sessions.

    Accessed Sep 28, 2026

  11. Claude Fable 5.1 and Mythos 5.1AnthropicOfficial

    CursorBench 3.2.0 and Terminal-Bench 4.0 score and cost at each effort level for Fable 5.1, Fable 5 and Mythos 5.1.

    Published Sep 1, 2026 · Accessed Sep 28, 2026

  12. Claude Fable models on your planAnthropicOfficial

    Fable models use limits faster and are capped at 50% of the weekly limit on Max and Team Premium; credits only on Pro and Team Standard.

    Accessed Sep 28, 2026

  13. Claude Opus 5.5AnthropicOfficial

    Score and cost per task at each effort level for Opus 5.5, Fable 5.1 and Opus 5 on Terminal-Bench 4.0, FrontierCode v1.1 and CursorBench 4.0; five-hour limit change of 2026-09-22.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  14. Claude Opus 5.5 System CardAnthropicOfficial

    Headline coding scores (SWE-bench Pro, Terminal-Bench 4.0 at xhigh and max, DeepSWE v1.1, CursorBench 4.0) versus Opus 5, Fable 5.1 and GPT-6 Astra.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  15. Claude pricingAnthropicOfficial

    Prices of Free, Pro, Max 5x, Max 20x, Team and Enterprise plans and what each includes.

    Accessed Sep 28, 2026

  16. Claude Sonnet 5 System CardAnthropicOfficial

    Headline coding scores of Sonnet 5 at max effort (SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.1).

    Published Jun 30, 2026 · Accessed Sep 28, 2026

  17. EffortAnthropicOfficial

    Effort levels available per model and when to use each (xhigh for agentic coding over 30 minutes, low for subagents).

    Accessed Sep 28, 2026

  18. How AI is transforming work at AnthropicAnthropicOfficial

    Claude Code transcripts by task: feature implementation 36.9%, design and planning 9.9%; daily use for debugging 55%.

    Published Dec 2, 2025 · Accessed Sep 28, 2026

  19. How do usage and length limits work?AnthropicOfficial

    Five-hour session and weekly limits; no fixed message count; what drives consumption (model, effort, length, tools).

    Accessed Sep 28, 2026

  20. Introducing Claude Haiku 4.5AnthropicOfficial

    SWE-bench Verified score of Haiku 4.5.

    Published Oct 15, 2025 · Accessed Sep 28, 2026

  21. Introducing Claude Sonnet 5AnthropicOfficial

    Release date, positioning and headline coding scores of Sonnet 5.

    Published Jun 30, 2026 · Accessed Sep 28, 2026

  22. Model configurationAnthropicOfficial

    Default model (Opus 5.5 on every plan since v2.1.280) and default effort per model in Claude Code.

    Accessed Sep 28, 2026

  23. Models overviewAnthropicOfficial

    Model IDs, release dates, context windows, max output and status of every current Claude model.

    Accessed Sep 28, 2026

  24. Models, usage, and limits in Claude CodeAnthropicOfficial

    Opus uses meaningfully more quota than Sonnet; higher effort reaches limits faster.

    Accessed Sep 28, 2026

  25. Optimizing for cost and intelligenceAnthropicOfficial

    Cost per solved task on a 478-problem SWE-bench Pro subset for Opus 5.5, Fable 5.1 and Sonnet 5 configurations; effort trade-offs for Opus 5.5.

    Accessed Sep 28, 2026

  26. PricingAnthropicOfficial

    Input, cache write, cache read and output prices per million tokens; batch and fast-mode prices.

    Accessed Sep 28, 2026

  27. Prompting Claude Opus 5.5AnthropicOfficial

    Start at medium; reserve xhigh and max for work with a measured gain; name the finish line; name the frontend styles to avoid.

    Accessed Sep 28, 2026

  28. What is the Max plan?AnthropicOfficial

    Max 5x and Max 20x usage relative to Pro per five-hour session; monthly billing only.

    Accessed Sep 28, 2026

  29. 5-hour session limits increaseAnthropic (ClaudeDevs on X)Official

    Five-hour limits up 20% on 2026-09-22; Opus 5.5 goes 25% further within limits because it is priced lower.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  30. Arena Code: WebDev leaderboardArena (LMArena)Third party

    Rating, 95% confidence interval, votes and rank of every model on blind pairwise web-development votes (entries[] in the server-rendered payload; vote cutoff 2026-09-25).

    Published Sep 25, 2026 · Accessed Sep 28, 2026

  31. Artificial Analysis Coding Agent IndexArtificial AnalysisThird party

    Coding Agent Index v1.5 per harness × model (Claude Code, Codex, Grok Build, Kimi Code, Antigravity, OpenCode): score, cost, tokens and time per task.

    Accessed Sep 28, 2026

  32. Artificial Analysis evaluation detailsArtificial AnalysisThird party

    Per-model sub-scores embedded in the page payload: Humanity's Last Exam, CritPt, SciCode, GPQA and Terminal-Bench 4.0 for each model × effort.

    Accessed Sep 28, 2026

  33. Artificial Analysis model leaderboardArtificial AnalysisThird party

    Intelligence Index v4.3 and AA's own Terminal-Bench 4.0 run for every model × effort, with cost per task and output tokens per task (JSON embedded in the page).

    Accessed Sep 28, 2026

  34. Grok 4.7 usage on SuperGrok HeavyBigGo Finance (reporting a user post)Third party

    About $40 of Grok 4.7 usage in Grok Build burned about 8% of the Heavy weekly limit.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  35. How I use Claude Code (thread)Boris Cherny (X)Third party

    Personal workflow of Claude Code's creator: give Claude a way to verify its work (his estimate: 2–3× the quality), shared CLAUDE.md, subagents, hooks. Read through a mirror (x.com answers 402).

    Published Jan 2, 2026 · Accessed Sep 28, 2026

  36. Claude Max weekly limit measurementsClassmethodThird party

    Weekly budget of Max 5x and Max 20x in API-equivalent dollars (two accounts, /usage vs ccusage).

    Published Aug 20, 2026 · Accessed Sep 28, 2026

  37. Codex usage researchcodexusage.devThird party

    Codex Plus 5-hour and weekly budgets in credits and dollars, cost per busy hour by effort level.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  38. Coding plan trackercodingplan.orgThird party

    Current Kimi plan names and CNY prices (Go, Plus, Pro, Max) and the mapping from legacy names.

    Published Sep 24, 2026 · Accessed Sep 28, 2026

  39. Devin pricingCognitionOfficial

    Devin Pro and Max prices.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  40. FrontierCode leaderboard dataCognitionThird party

    FrontierCode v1.1 score (new_score), cost, tokens and duration for every model × effort, main and extended subsets.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  41. Cursor pricingCursorOfficial

    Cursor Pro, Pro+ and Ultra prices and their two usage pools.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  42. CursorBench 4.0CursorThird party

    Score, cost per task, tokens and steps for every model × effort (server-rendered results table).

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  43. DeepSWE v1.1 leaderboardDatacurveThird party

    pass@1 and mean cost per task for every model × effort (mini-swe-agent harness, 4 runs).

    Accessed Sep 28, 2026

  44. GLM-5.3 in Claude Code: real cost on the Coding PlanDécodeur IAThird party

    Lite plan in Claude Code: a 3-hour session used 70% of a 2,000-credit window; one endpoint with tests took 18 minutes and 210 credits.

    Published Aug 29, 2026 · Accessed Sep 28, 2026

  45. Design Arena: agentic web devDesign ArenaThird party

    Elo, standard error, win rate and battles for the full-stack and frontend agentic web-app arenas (embedded boards and /api/leaderboard).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  46. Design Arena: data visualizationDesign ArenaThird party

    Elo, battles and win rate of 163 models on blind human votes for chart and dashboard code (POST /api/leaderboard, category dataviz, 79,202 votes; no standard error published).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  47. Design Arena: SVGDesign ArenaThird party

    Elo, standard error, battles and win rate of 123 models on blind human votes for SVG generation (POST /api/leaderboard, category svg, 190,054 votes).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  48. Factory pricingFactoryOfficial

    Droid Pro, Plus and Max prices and per-model multipliers.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  49. Kimi plans in USDgeotoolbox.aiThird party

    International USD prices of Kimi plans and their Kimi Code usage multipliers (1×, 5×, 15×, 30×).

    Published Sep 5, 2026 · Accessed Sep 28, 2026

  50. GitHub Copilot plansGitHubOfficial

    Copilot Pro, Pro+, Max, Business and Enterprise prices and the AI Credits included with each.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  51. Pro 20x weekly meter measurementGitHub user reportThird party

    Weekly meter moved from 64% to 95% for $609 of API-equivalent usage on Pro 20x.

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  52. Three weekly meters on Pro 5xGitHub user reportThird party

    Three full weekly Pro 5x meters = $959.87 API-equivalent (about $320 per week) and 708 M tokens on GPT-5.6 Sol and Astra.

    Published Sep 7, 2026 · Accessed Sep 28, 2026

  53. Weekly allowance value on Pro 20x (GPT-5.6 Sol vs GPT-6 Astra)GitHub user reportThird party

    API-equivalent value of a full weekly Pro 20x allowance: about $1,925 on GPT-5.6 Sol, $1,191–1,234 on GPT-6 Astra.

    Published Sep 8, 2026 · Accessed Sep 28, 2026

  54. Claude Code usage meter measurementsGitHub user reportsThird party

    Controlled A/B measurements of API-equivalent dollars per 1% of the 5-hour and weekly meters (interactive vs headless), Pro and Max.

    Published Sep 27, 2026 · Accessed Sep 28, 2026

  55. Antigravity CLI best practicesGoogleOfficial

    Verification loops first; explore, plan, then execute; `/rewind`, `/fork`, parallel subagents.

    Accessed Sep 28, 2026

  56. Antigravity modelsGoogleOfficial

    Models available per plan and thinking options (Low/Medium/High) in Antigravity; two usage pools with five-hour and weekly remaining percentages.

    Accessed Sep 28, 2026

  57. Antigravity plansGoogleOfficial

    Quota wording per plan: weekly refresh on Free, five-hour refresh until a weekly limit on Pro and Ultra.

    Accessed Sep 28, 2026

  58. Changes to Antigravity plansGoogleOfficial

    Ultra $100 = 5× and $200 = 20× the Pro token allowance; the Gemini pool is drawn down at API prices; Claude and GPT models use a separate fixed pool.

    Published May 19, 2026 · Accessed Sep 28, 2026

  59. Gemini API pricingGoogleOfficial

    Prices per million tokens for Gemini 3.8/3.7 Flash (introductory until 2026-12-31) and 3.1 Pro Preview.

    Published Sep 24, 2026 · Accessed Sep 28, 2026

  60. Gemini thinkingGoogleOfficial

    Thinking levels and defaults per Gemini model; guidance per level.

    Published Sep 25, 2026 · Accessed Sep 28, 2026

  61. Google AI subscriptionsGoogleOfficial

    Prices of Google AI Plus, Pro, Ultra 5x and Ultra 20x; Gemini app usage multipliers.

    Accessed Sep 28, 2026

  62. Gemini 3.1 Pro model cardGoogle DeepMindOfficial

    SWE-bench Verified, SWE-bench Pro and Terminal-Bench 2.0 scores of Gemini 3.1 Pro (high thinking).

    Published Feb 19, 2026 · Accessed Sep 28, 2026

  63. Gemini 3.8 Flash evaluationGoogle DeepMindOfficial

    Coding benchmark table of Gemini 3.8 Flash versus Opus 5, Sonnet 5 and GPT-5.6, with the settings used.

    Published Sep 2, 2026 · Accessed Sep 28, 2026

  64. Boris Cherny on plan mode (Hacker News)Hacker NewsThird party

    “plan mode was useful, and is no longer useful” with Opus 5.5 and Fable; `/plan` stays available (HN Algolia API).

    Published Sep 25, 2026 · Accessed Sep 28, 2026

  65. SuperGrok Grok Build usage reportHacker NewsThird party

    Four terminals at full speed for one hour used 9% of the SuperGrok weekly pool during a 2× promotion.

    Published Aug 12, 2026 · Accessed Sep 28, 2026

  66. Qwen3.8-27B hardware testsHardware CornerThird party

    Local decode speeds of Qwen3.8-27B on RTX 4090, 3090, 5090 and M5 Max at several context lengths.

    Published Aug 17, 2026 · Accessed Sep 28, 2026

  67. AI coding agent adoption 2026JetBrains ResearchThird party

    Share of developers using each coding agent (Claude Code 39%, Copilot 21%, Codex 16%, Cursor 12%, OpenCode 7%, Antigravity 6%; 15,000+ respondents).

    Published Aug 15, 2026 · Accessed Sep 28, 2026

  68. KiloBenchKiloThird party

    Terminal-Bench 2.0 completion and cost per full attempt in the Kilo agent (embedded chart points).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  69. LiveBenchLiveBenchThird party

    Coding and agentic-coding category scores and cost per question (table_2026_06_25.csv and category files).

    Published Sep 10, 2026 · Accessed Sep 28, 2026

  70. Best AI for codingllm-stats (ZeroEval)Third party

    TrueSkill coding rating (μ − 3σ) over 60 benchmark games for every model (initialIndexData in the payload). Needs a cache-busting query string: the plain URL served a stale 2026-09-22 copy.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  71. ZeroEval text-to-SVG arenallm-stats (ZeroEval)Third party

    Conservative TrueSkill ratings (1,142 votes); last updated 2026-08-05, without Opus 5.5, Fable 5.1 or any GPT-6 model: kept only to show it is stale.

    Published Aug 5, 2026 · Accessed Sep 28, 2026

  72. Mac Studio M5 Ultra reviewMacStoriesThird party

    Local decode speed of GLM-5.3-Flash on the M5 Ultra.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  73. Muse CodeMetaOfficial

    Muse Code plans (Everyday, High, Power) and prompts per 5 hours.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  74. MiniMax Token PlanMiniMaxOfficial

    Token Plan prices and 5-hour / weekly windows.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  75. Mistral pricingMistral AIOfficial

    Vibe Pro and Team prices (no published quota).

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  76. Kimi API pricingMoonshot AIOfficial

    Cache hit, cache miss and output prices per million tokens for Kimi K3 and K2.7 Code.

    Accessed Sep 28, 2026

  77. Kimi Code membershipMoonshot AIOfficial

    Five-hour rolling window, monthly cap shared across Kimi services, extra usage top-ups.

    Accessed Sep 28, 2026

  78. Kimi Code modelsMoonshot AIOfficial

    Models in Kimi Code (K3, K2.8 Preview, HighSpeed at 3× quota), effort levels and plan access.

    Accessed Sep 28, 2026

  79. Kimi K3Moonshot AIOfficial

    Coding benchmark table of Kimi K3 (max effort) versus Fable 5, GPT-5.6 Sol, Opus 4.8 and GLM-5.2, with harnesses used.

    Published Jul 16, 2026 · Accessed Sep 28, 2026

  80. Kimi membership pricingMoonshot AIOfficial

    CNY monthly prices of the membership tiers (legacy names).

    Accessed Sep 28, 2026

  81. AIDev-pop: pull requests opened by coding agentsMSR 2026 (arXiv)Third party

    Types of 33,596 agent pull requests: features 43.0%, fixes 24.1%, docs 11.6%, tests 7.0%, refactors 6.8%.

    Published Feb 3, 2026 · Accessed Sep 28, 2026

  82. SWE-rebenchNebiusThird party

    Resolved rate, pass@5 and cost per task on the 2026-05-15 → 2026-07-01 window (embedded items).

    Published Jul 24, 2026 · Accessed Sep 28, 2026

  83. Cross-provider quota calibrationoh-my-pi (GitHub)Third party

    Tokens and API-equivalent dollars per 1% of the Anthropic, OpenAI and Antigravity meters, measured on one machine.

    Published Sep 26, 2026 · Accessed Sep 28, 2026

  84. About ChatGPT Pro tiersOpenAIOfficial

    Pro $100 (5× Plus) and Pro $200 (20× Plus); new Pro $200 sign-ups paused since 2026-09-10.

    Accessed Sep 28, 2026

  85. API pricingOpenAIOfficial

    Standard, batch, flex and fast-mode prices per million tokens for GPT-6 Astra, Sol and Luna.

    Accessed Sep 28, 2026

  86. ChatGPT rate card (credit-based pricing)OpenAIOfficial

    Credits per million tokens for each model (25× the API dollar price); fast mode 2.5×; Ultra billing.

    Accessed Sep 28, 2026

  87. Codex best practicesOpenAIOfficial

    Goal / Context / Constraints / Done when; starting efforts per model; `/plan`, interview and PLANS.md; update AGENTS.md after a repeated mistake; worktrees for parallel tasks.

    Accessed Sep 28, 2026

  88. Codex modelsOpenAIOfficial

    Effort controls per surface (Light to Ultra), recommended starting presets (Sol Medium, Luna High, Astra Light), surfaces per model.

    Accessed Sep 28, 2026

  89. Codex pricingOpenAIOfficial

    Plan prices and the estimated local messages per 5 hours for each model on Plus, Pro 5x, Pro 20x and Business.

    Accessed Sep 28, 2026

  90. GPT-6 AstraOpenAIOfficial

    Terminal-Bench 4.0 score, cost, output tokens and latency at each effort for GPT-6 Astra, GPT-5.6 Sol and Fable 5.1; FrontierCode 1.1 Extended; output tokens per DeepSWE task.

    Published Sep 3, 2026 · Accessed Sep 28, 2026

  91. GPT-6 Astra model page (and /gpt-6-sol, /gpt-6-luna)OpenAIOfficial

    Release dates, context windows, knowledge cutoffs and supported effort values.

    Accessed Sep 28, 2026

  92. Introducing GPT-6 Sol and LunaOpenAIOfficial

    DeepSWE v1.1 and FrontierCode 1.1 Main score and cost per task at each effort for GPT-6 Astra, Sol, Luna, GPT-5.6, Opus 5 and Fable 5.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  93. Managing billing and seats in ChatGPT BusinessOpenAIOfficial

    Business Standard and Premium seat prices, annual and monthly.

    Accessed Sep 28, 2026

  94. Managing usage with GPT-6 Astra in Work and CodexOpenAIOfficial

    Five-hour and weekly windows, shared allowance across Codex and Work, effort advice (Astra Low can beat Sol High).

    Accessed Sep 28, 2026

  95. OpenCode GoOpenCodeOfficial

    OpenCode Go price and requests per 5 hours for each open-weight model.

    Published Sep 28, 2026 · Accessed Sep 28, 2026

  96. OpenRouter rankings: top models by taskOpenRouterThird party

    Share of spend and tokens per task over 30 days (code: general implementation 10.1%, debugging 5.2%, file edits 4.3%, review 2.8%, front-end UI 2.0%; math 0.8%), from /api/frontend/v1/rankings/task-spend.

    Published Sep 27, 2026 · Accessed Sep 28, 2026

  97. FrontierSWE V2Proximal LabsThird party

    Mean, best and worst of 5 trials, cost per trial and duration on 34 long-horizon tasks.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  98. Qwen3.8-27B quantizations benchmarkedQuesmaThird party

    Quality of Qwen3.8-27B by quantization level on Terminal-Bench 2.1 (Q4_K_M matches BF16, Q2 loses about 10 points).

    Published Aug 26, 2026 · Accessed Sep 28, 2026

  99. Terminal-Bench 4.0 leaderboardTerminal-Bench (tbench.ai)Third party

    Accuracy with confidence interval, total cost and tokens per agent × model × effort over 330 trials.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  100. Vals AI benchmarksVals AIThird party

    Accuracy, cost per test and latency for Terminal-Bench 4, Vibe Code Bench, Vals Index and other coding benchmarks (astro-island props).

    Published Sep 23, 2026 · Accessed Sep 28, 2026

  101. Vals IndexVals AIThird party

    Vals Index score, cost per test and latency for Opus 5.5, Fable 5.1, GPT-6 Astra and others.

    Accessed Sep 28, 2026

  102. Grok 4.6xAIOfficial

    Coding scores of Grok 4.6 at high effort (DeepSWE v1.1, Terminal-Bench 3.0, CursorBench 3.2).

    Published Aug 12, 2026 · Accessed Sep 28, 2026

  103. Grok 4.7xAIOfficial

    CursorBench 4.0 score, cost per task, output tokens and steps at each effort for Grok 4.7 and for Fable 5.1, Opus 5, Sonnet 5, GPT-5.6 Sol/Terra/Luna and Gemini 3.8 Flash (JS dataset behind the chart).

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  104. Grok 4.7 model cardxAIOfficial

    Terminal-Bench 4.0, FrontierSWE V2 and SWE-Marathon scores of Grok 4.7 and 4.6; CursorBench 4.0 chart (vector geometry) for Grok 4.6.

    Published Sep 21, 2026 · Accessed Sep 28, 2026

  105. Grok Build plan modexAIOfficial

    When to plan (architecture ambiguity) and when to skip it (obvious change).

    Accessed Sep 28, 2026

  106. Grok FAQxAIOfficial

    One shared weekly usage pool across Chat, Imagine, Voice and Build since June 2026; shown only as a percentage; extra usage credits.

    Published Aug 27, 2026 · Accessed Sep 28, 2026

  107. Grok plansxAIOfficial

    Monthly and annual prices of SuperGrok Lite, SuperGrok, Plus and Heavy, from the product JSON the page loads (grok.com/rest/products).

    Accessed Sep 28, 2026

  108. ReasoningxAIOfficial

    Effort levels per Grok model (low to xhigh, default high) and guidance per level.

    Accessed Sep 28, 2026

  109. xAI API pricingxAIOfficial

    Input, cached input and output prices per million tokens for Grok 4.7, 4.6 and grok-build-0.1.

    Accessed Sep 28, 2026

  110. MiMo Token PlanXiaomiOfficial

    MiMo Token Plan prices and credits.

    Published Sep 22, 2026 · Accessed Sep 28, 2026

  111. GLM Coding PlanZ.aiOfficial

    Prices of Lite, Pro and Max (monthly, quarterly, yearly) and their five-hour and weekly credit allowances.

    Accessed Sep 28, 2026

  112. GLM Coding Plan overviewZ.aiOfficial

    Credit formula and per-model token multipliers, peak and off-peak rates, and the official weekly token estimates per tier.

    Accessed Sep 28, 2026

  113. GLM-5.3 launch postZ.aiOfficial

    Z.ai Code Bench accuracy and output tokens per task at each effort for GLM-5.3, Fable 5 and Opus 4.8.

    Published Aug 14, 2026 · Accessed Sep 28, 2026

  114. GLM-5.3 model cardZ.aiOfficial

    Coding benchmark table of GLM-5.3 at max effort in Claude Code (Terminal-Bench 2.1 and 3.0, DeepSWE v1.1, NL2Repo, FrontierSWE).

    Accessed Sep 28, 2026

  115. GLM-5.3 model guideZ.aiOfficial

    Effort levels (low, high, max; default max) and the recommendation to use max for coding.

    Accessed Sep 28, 2026

  116. GLM-5.3-FlashZ.aiOfficial

    Coding scores of GLM-5.3-Flash (Terminal-Bench 2.1, DeepSWE v1.1) from printed chart labels.

    Published Aug 26, 2026 · Accessed Sep 28, 2026

  117. Z.ai API pricingZ.aiOfficial

    Input, cached input and output prices per million tokens for GLM-5.3 and GLM-5.3-Flash.

    Accessed Sep 28, 2026