Lena is here. I went through the official docs and migration materials after the April 16 release, expecting most of it to be familiar. It mostly is. But there are enough structural changes — three breaking API changes, a new effort system, a tokenizer that shifts token counts — that treating this as a simple model ID swap would be a mistake.
Here's what I found when I worked through it.
What the Claude Opus 4.7 API Includes
Model ID, Context Window, Output Limits, and Tools
The basics first, because they're worth stating clearly.
**Model ID: claude-opus-4-7**. According to Anthropic's official models overview, it supports a 1M token context window at standard API pricing with no long-context premium, and 128k max output tokens on the synchronous Messages API. On the Message Batches API, Opus 4.7 can go up to 300k output tokens using the output-300k-2026-03-24 beta header.
The full tool set carries over from Opus 4.6: bash, code execution, computer use, text editor, web search, web fetch, MCP connector, and memory tools are all available on day one. Vision support is present across the board — and meaningfully upgraded, which I'll come back to.
Available across the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry — and rolling out on GitHub Copilot for Copilot Pro+, Business, and Enterprise users. Pricing: $5 per million input tokens, $25 per million output tokens — unchanged from Opus 4.6.
What Is New vs Opus 4.6
Three things are genuinely new in the API surface:
Adaptive thinking as the only thinking mode. The old {"type": "enabled", "budget_tokens": N} pattern is gone. Sending it now returns a 400 error. Opus 4.7 uses {"type": "adaptive"} — the model decides dynamically how much to reason based on task complexity. Adaptive thinking is off by default; you must enable it explicitly if you want the model to think at all.
The xhigh effort level. This sits between high and max, and it's the new recommended starting point for coding and agentic use cases. The API default is high; you set xhigh explicitly through output_config. I'll show the structure below.
Task budgets (beta). A new mechanism that gives the model an advisory token target for an entire agentic loop — thinking, tool calls, tool results, and final output combined. The model sees a running countdown and uses it to prioritize work and wrap up gracefully as the budget runs out.
Also removed: non-default sampling parameters. Setting temperature, top_p, or top_k to any non-default value now returns a 400 error. If you were using temperature=0 for determinism, note that it never guaranteed identical outputs anyway — the migration guide recommends omitting these parameters entirely.
A Minimal Claude Opus 4.7 API Setup
First Request Structure
The minimum working request for Opus 4.7 looks like this:
import anthropic
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-opus-4-7",
max_tokens=4096,
messages=[
{"role": "user", "content": "Explain the tradeoffs between BFS and DFS for a graph with cycles."}
]
)
print(message.content[0].text)
This runs without thinking enabled. The model will respond directly. For most tasks, this is the right starting point — add complexity only when you have a reason to.
Choosing Adaptive Thinking and Effort Levels
When you want the model to reason before responding, add the thinking configuration and set effort explicitly:
message = client.messages.create(
model="claude-opus-4-7",
max_tokens=16384,
thinking={"type": "adaptive"},
output_config={"effort": "xhigh"},
messages=[
{"role": "user", "content": "Review this pull request for security vulnerabilities..."}
]
)
A few things worth knowing about effort levels, based on Anthropic's official effort documentation:
highis the API default. Use it for complex reasoning, nuanced analysis, or difficult coding problems where quality is the priority.xhighis the new level — recommended for coding and agentic tasks. Claude Code has raised its default to xhigh across all plans.maxprovides the deepest reasoning with no token constraint. It applies to the current session only (unless set via environment variable) and doesn't persist.low and **medium** trade accuracy for speed and cost. Useful for high-volume classification or routing where marginal quality differences don't justify spending.
One detail I went back and re-read twice: Opus 4.7 respects effort levels more strictly than Opus 4.6, especially at low and medium. If you observe shallow reasoning on a complex task, the right move is to raise effort — not to add prompt scaffolding around it. The docs are explicit about this.
When running at xhigh or max, set max_tokens to at least 64k to give the model room to think and act across subagents and tool calls. Starting at 64k and tuning from there is Anthropic's own recommendation.
Features That Matter for Long-Running Agents
High-Resolution Vision, xhigh Effort, and Tool Workflows
The vision upgrade is the most concrete capability gain for agent builders. As documented in Vellum AI's benchmark analysis of Opus 4.7, OSWorld-Verified (computer use) climbed from 72.7% on Opus 4.6 to 78.0% — a 5-point gain that, paired with the resolution upgrade, shifts the economics of UI automation meaningfully.
Opus 4.7 is the first Claude model with high-resolution image support: maximum resolution increased from 1,568 pixels (~1.15MP) to 2,576 pixels (~3.75MP) on the long edge. That's more than triple the pixel budget. For computer-use agents that read dense UIs, screenshot-based workflows, or document understanding pipelines, this is a meaningful change. Critically, coordinates now map 1:1 with actual image pixels — the scale-factor math that was previously required for coordinate extraction is gone.
Vision in a request:
message = client.messages.create(
model="claude-opus-4-7",
max_tokens=4096,
messages=[
{
"role": "user",
"content": [
{
"type": "image",
"source": {"type": "url", "url": "https://example.com/diagram.png"}
},
{"type": "text", "text": "List every service shown and the connections between them."}
]
}
]
)
If the extra resolution isn't needed for a given task, downsample images before sending — high-resolution images produce more tokens, and for cost-sensitive workloads that adds up.
For tool-heavy agentic loops, raising effort increases the frequency and depth of tool calls. The relationship is direct: lower effort → fewer tool calls and shallower reasoning chains; higher effort → more thorough tool interaction. This is steerable through prompting as well, but the effort parameter is the cleaner lever.
Cost and Latency Controls for Production Use
Task budgets are the new mechanism for controlling spend on long agentic loops. Enable them with the beta header:
response = client.beta.messages.create(
model="claude-opus-4-7",
max_tokens=128000,
output_config={
"effort": "high",
"task_budget": {"type": "tokens", "total": 128000}
},
messages=[
{"role": "user", "content": "Review the codebase and propose a refactor plan."}
],
betas=["task-budgets-2026-03-13"]
)
The model sees the countdown and uses it to prioritize and wrap up gracefully. Without a task budget, default behavior is "spend as needed" — on a complex agentic task at xhigh, that can mean significantly more output tokens than you'd expect from a single-turn request.
For async workloads — evaluation runs, nightly summaries, batch analysis — the Batch API gives a 50% discount and removes rate-limit pressure from real-time traffic. Prompt caching remains available and can cut repeated input costs by up to 90% for workloads with stable system prompts or large static prefixes.
Migration Mistakes to Avoid
Assuming 4.6 Prompts Transfer Unchanged
This is the one I keep seeing come up.
Opus 4.7 follows instructions more literally than Opus 4.6. It no longer reads between the lines or silently generalizes from one case to another. Soft phrasings — "try to," "if possible," "roughly" — now carry more actual weight. Prompts that relied on 4.6's interpretive flexibility will sometimes behave differently, and not always in the way you'd expect.
The three breaking API changes require code updates, not just prompt edits:
- Replace
thinking: {type: "enabled", budget_tokens: N}withthinking: {type: "adaptive"} - Remove
temperature,top_p,top_kfrom requests entirely - Audit
max_tokens— the new tokenizer maps the same text to up to 1.35× more tokens, so existing ceilings may cut off responses that previously fit
On tone: Opus 4.7 is more direct and opinionated than 4.6 — fewer emoji, less validation-forward phrasing. If your product relies on a specific voice calibrated to 4.6's warmer style, re-evaluate your style prompts against the new baseline before promoting to production.
Anthropic's Opus 4.7 launch announcement links directly to the full migration checklist, and Claude Code users can run /claude-api migrate to automate the model ID swap and breaking parameter changes across a codebase.
Treating Long Context Like Persistent Memory
I paused here when I read this in the docs, because it's easy to conflate the two.
A 1M token context window is not persistent memory. It means you can fit 1 million tokens into a single request. When that request ends — when the session closes, when the agent crashes, when a new conversation starts — that context is gone. The next request starts at zero.
Opus 4.7 does include improved file system-based memory: the model reads and writes to notes files across multi-session work, with noticeably more reliable behavior for agents using this pattern. But that's a tool you configure. It doesn't happen automatically from a long context window.
The distinction matters for anyone building agents that are supposed to "remember" things across runs. The 1M window helps within a session. Cross-session memory requires explicit architecture.
What the API Still Does Not Solve
Capability Reuse and Validated Repair History
This is the part I find harder to fit neatly into an API guide, but I think it's worth saying.
When Opus 4.7 catches a logical fault during the planning phase — which the release materials describe as a genuine capability — that reasoning happens within the session. The fact that it caught the fault, and how it corrected it, doesn't automatically persist as a reusable pattern for the next run. The next time the same class of problem appears, the model reasons from scratch.
According to a research paper on AI agent reliability published in early 2026, most models are benchmarked on average accuracy rather than consistency across runs — which means a model can score well on benchmarks while still failing unpredictably on the same class of task at different times. Opus 4.7 improves on Opus 4.6's baseline. That's real. But in-session correction and cross-session capability inheritance are different problems, and the API solves the first, not the second.
For builders running production agents, this means the engineering work of building reliable systems — evaluation harnesses, repair documentation, monitoring — still sits outside the model itself. The API gives you a more capable model. What you do with that capability over time is still on your architecture.
FAQ
Q: What is the correct model ID for Opus 4.7?
A: claude-opus-4-7. This is the stable model string for API calls across the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
Q: Is adaptive thinking on by default?
A: No. Adaptive thinking is off by default on Opus 4.7. You must set thinking: {"type": "adaptive"} explicitly to enable it. Requests with no thinking field run without thinking.
Q: What happens if I send temperature or budget_tokens?
A: Both return a 400 error on Opus 4.7. Remove temperature, top_p, and top_k from all requests. Replace budget_tokens with output_config: {"effort": "..."} and thinking: {"type": "adaptive"}.
Q: When should I use xhigh vs high effort?
A: Use xhigh as the starting point for coding and agentic tasks — it's now the default in Claude Code across all plans. Use high for most intelligence-sensitive tasks. Drop to medium or low when latency or cost is more important than reasoning depth. If you observe shallow output on a complex task at a lower level, raise effort rather than adding prompt scaffolding.
Q: Does the 1M context window mean the agent remembers things between sessions?
A: No. The context window applies within a single request. When a session ends, that context is gone. Multi-session memory requires explicit tooling — file-based memory, external stores, or similar architecture.
Previous Posts
- 👉 If you want to understand why model upgrades don’t actually reduce system cost:Claude Opus 4.7 vs Reliability: Why Better Models Don’t Fix Agent Systems
- 👉 If you're trying to understand where real cost hides beyond token pricing:Harness Engineering: The Hidden Layer Behind Agent Cost and Reliability
- 👉 If you're evaluating when Opus is actually worth the premium over Sonnet:Claude Managed Agents: When You Should Actually Use Opus vs Sonnet
- 👉 If you're building agent systems and need to control behavior + spend together:Agent Superpowers: Behavior Constraints as a Cost Control Layer
- 👉 If you're thinking about scaling agent usage beyond single workflows:Best MCP Servers for Claude Code: Real Production Use Cases




