# Claude Code effort levels Effort tells Claude roughly how much compute you want it to spend on a task. In [[Claude Code]] there are five levels (`low`, `medium`, `high`, `xhigh`, `max`), and you can change it at any time with `/effort`. On [[Claude Opus 5.5]] and [[Claude Fable 5.1]], switching mid-conversation doesn't break the [[Claude Code Prompt Caching|prompt cache]], so moving up and down within a session costs nothing extra. [[Thariq Shihipar]] (Claude Code team at [[Anthropic]]) ran the same builds at every level and went through the Terminal-Bench 3.0 results task by task. His finding: higher effort mostly buys more checking. It changes how much verification, edge-case testing and independent judgement Claude does, and it does very little for a task Claude approached the wrong way. ## Think of it as a deadline That's his analogy. Give someone 12 hours for a task and they'll assume you want them to try very hard. Give the same person 1 hour and they'll hand you the best version that fits, expecting to iterate from there. Or they push back, say it needs at least 3 hours, and take 3. Claude reads effort the same way. It always tries to do the task reasonably, but at higher effort it does more on its own (more decisions, more tests, more verification) before coming back to you. You pay for that in tokens. On Terminal-Bench 3.0, the median Fable 5.1 attempt used about 73k tokens at low and about 222k at max. On the newest models each step up adds both score and tokens, which makes the setting predictable. ## The less you specify, the more effort decides for you Three experiments on Opus 5.5: - **Underspecified**: "build a personal fitness and workout tracker app". Low (1.5 min) gave a log and a simple graph. Max (67 min) gave a much richer app, heat chart included, and made more choices on his behalf - **Lightly specified**: redesign the `/config` menu. Every level landed on the same idea (submenus, better search). Low (1 min) produced an interactive sketch that didn't look much like Claude Code; max (28 min) produced a faithful mockup with walkthroughs of several flows. He prefers low here, to see Claude's vision quickly and give feedback - **Highly specified**: Claude interviewed him first, then implemented that spec. All levels converged on similar designs and implementations; max mostly spent its extra time simplifying details For feature work, then, picking a level mostly means picking how much you want to [[Human-in-the-Loop|stay in the loop]]. A detailed spec shrinks the gap between levels, so the spec (i.e., good [[Context Engineering]]) is the cheaper lever. ## Effort buys edge-case coverage His main takeaway from the benchmark: "higher effort is best for tasks with lots of hidden edge cases." The cleanest example is `html-js-filter`, an HTML sanitizer that must strip every way of smuggling JavaScript into a page. Fable 5.1 went from 1/5 at low to 5/5 at xhigh. A low attempt took about 2 minutes: it wrote the filter in one pass and tested it against one hand-written page. A high-effort run took about 33 minutes. It reviewed its own draft adversarially, read the installed parser's source to look for bugs, checked that clean documents came out unchanged, ran a standard XSS test suite, and finally wrote a random-document fuzzer. The failure breakdown across all 370 Fable 5.1 attempts shows where the gain comes from: | Outcome (Fable 5.1) | Low | Max | |---|---|---| | Passed | 140 | 214 | | Missed a case | 59 | 24 | | ...of which a bug its tests missed | 40 | 14 | | Made the wrong call | 133 | 107 | | ...of which picked the wrong reading of the task | 25 | 47 | Missed cases drop by more than half. Wrong calls barely move, and misreading the task got MORE frequent at max. (Anthropic notes that a model judge labeled the failure kinds, so treat them as approximate.) Three Opus 5.5 tasks that failed at low and passed at higher effort show the same habits: reproduce first, test against a reference, try to break your own work. - **`mvcc-lsm-compaction`** (fix a storage-engine bug from its crash report, 0/5 → 4/5 at xhigh): at low, Claude edited code before building or running the reproducer. At xhigh (about 11 minutes instead of 1), it reproduced the crash first, wrote a randomized test against a reference that never compacts, and checked that its tests failed on half-finished fixes - **`cli-2ph-simplex`** (a linear-program solver CLI in Python, 0/5 → 5/5 at high): low stopped around 10k tokens and warned it might be slow on big problems without checking. High compared it against a brute-force solver on random problems, timed bigger ones, hit crashes and timeouts, and reworked the search - **`gsea-proteomics`** (gene set enrichment analysis on proteomics data, 0/5 → 4/5 at high): low picked one reasonable way to prep the data and reported the result. High tried two, noticed the significant treatments changed, and dug into why before choosing. With a user in the loop, Claude might simply have asked ## Where it pays off most Fable 5.1 pass rate, low effort → top effort, per Terminal-Bench 3.0 category: | Category | Low | Top | |---|---|---| | Security (7 tasks) | 64% | 87% | | Hardware (5 tasks) | 34% | 75% | | ML (13 tasks) | 54% | 73% | | Science (15 tasks) | 41% | 61% | | Software (20 tasks) | 43% | 56% | | Media (4 tasks) | 18% | 30% | | Operations (10 tasks) | 12% | 22% | "Low" pools each model's two lowest settings and "top" its three highest, because the categories are small. Thariq names hardware, code review and security as the big winners, domains where verification finds real bugs. Rulebook-style work (operations, like running a month-end EU trade-statistics filing) stays low at every level: extra tokens don't help when Claude doesn't know the rules. ## Which level to pick His rule of thumb: - **Low**: quick, in-the-loop work (brainstorming, sketching, easy changes) - **Medium**: most regular software engineering, such as new features. It's the default on Opus 5.5 - **High**: when verification matters or edge cases lurk, such as fixing a bug in a brownfield codebase - **Max**: Claude working fully autonomously on hard problems, such as building and verifying an app end to end, or hunting security vulnerabilities in critical software And the loop he uses for feature work: 1. Give Claude a spec and ask it to interview you about the missing details 2. Implement on low 3. Review that it got the gist right; iterate on low 4. Verify and test on high ## My take Before picking a level, I'd ask what failure I'm trying to prevent. If it's a missed edge case, raise effort. If it's Claude solving the wrong problem, more tokens won't help; spend the time on the spec, or stay in the loop at low effort and steer. The numbers come from Anthropic's internal runs (5 attempts per task, production safety interventions off for Fable 5.1, security tasks without internet access), so they won't line up with the public leaderboard. And max has its own failure mode: in community tests it can overthink and burn its whole budget without an answer (see [[Claude Opus 5.5]]). I'd keep medium as the default. ## References - Thariq Shihipar, "Using Claude Code: Spending your effort", claude.dev, 2026-09-25: https://claude.dev/blog/spending-your-effort/ - Terminal-Bench 3.0 task list: https://github.com/harbor-framework/terminal-bench/releases/tag/v3.0.0 ## Related - [[Claude Code]] - [[Claude Opus 5.5]] - [[Claude Fable 5.1]] - [[Claude]] - [[Thariq Shihipar]] - [[Context Engineering]] - [[2026-10-04 Claude effort levels - Draft on low, verify on high]] - [[Claude Sonnet 5.5]] - [[Claude 5]] - [[AI Reasoning Models]] - [[Lydia Hallie]] - [[Claude Code Tips and Best Practices]]