EvoMap
How to Use the Claude Opus 4.7 API

How to Use the Claude Opus 4.7 API

April 24, 2026
206 views
Claude Opus 4.7 API adaptive thinking xhigh task budgets migration

Lena is here. I went through the official docs and migration materials after the April 16 release, expecting most of it to be familiar. It mostly is. But there are enough structural changes — three breaking API changes, a new effort system, a tokenizer that shifts token counts — that treating this as a simple model ID swap would be a mistake.

Here's what I found when I worked through it.

What the Claude Opus 4.7 API Includes

Model ID, Context Window, Output Limits, and Tools

The basics first, because they're worth stating clearly.

​**Model ID: ​claude-opus-4-7**​. According to Anthropic's official models overview, it supports a 1M token context window at standard API pricing with no long-context premium, and 128k max output tokens on the synchronous Messages API. On the Message Batches API, Opus 4.7 can go up to 300k output tokens using the output-300k-2026-03-24 beta header.

The full tool set carries over from Opus 4.6: bash, code execution, computer use, text editor, web search, web fetch, MCP connector, and memory tools are all available on day one. Vision support is present across the board — and meaningfully upgraded, which I'll come back to.

Available across the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry — and rolling out on GitHub Copilot for Copilot Pro+, Business, and Enterprise users. Pricing: $5 per million input tokens, $25 per million output tokens — unchanged from Opus 4.6.

What Is New vs Opus 4.6

Three things are genuinely new in the API surface:

Adaptive thinking as the only thinking mode. The old {"type": "enabled", "budget_tokens": N} pattern is gone. Sending it now returns a ​400 error​. Opus 4.7 uses {"type": "adaptive"} — the model decides dynamically how much to reason based on task complexity. Adaptive thinking is off by default; you must enable it explicitly if you want the model to think at all.

The ​xhigh​​ effort level. This sits between high and max, and it's the new recommended starting point for coding and agentic use cases. The API default is high; you set xhigh explicitly through output_config. I'll show the structure below.

Task budgets (beta). A new mechanism that gives the model an advisory token target for an entire agentic loop — thinking, tool calls, tool results, and final output combined. The model sees a running countdown and uses it to prioritize work and wrap up gracefully as the budget runs out.

Also removed: ​non-default sampling parameters​. Setting temperature, top_p, or top_k to any non-default value now returns a 400 error. If you were using temperature=0 for determinism, note that it never guaranteed identical outputs anyway — the migration guide recommends omitting these parameters entirely.

A Minimal Claude Opus 4.7 API Setup

First Request Structure

The minimum working request for Opus 4.7 looks like this:

Python
import anthropic

client = anthropic.Anthropic()

message = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=4096,
    messages=[
        {"role": "user", "content": "Explain the tradeoffs between BFS and DFS for a graph with cycles."}
    ]
)

print(message.content[0].text)

This runs without thinking enabled. The model will respond directly. For most tasks, this is the right starting point — add complexity only when you have a reason to.

Choosing Adaptive Thinking and Effort Levels

When you want the model to reason before responding, add the thinking configuration and set effort explicitly:

Python
message = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=16384,
    thinking={"type": "adaptive"},
    output_config={"effort": "xhigh"},
    messages=[
        {"role": "user", "content": "Review this pull request for security vulnerabilities..."}
    ]
)

A few things worth knowing about effort levels, based on Anthropic's official effort documentation:

  • high is the API default. Use it for complex reasoning, nuanced analysis, or difficult coding problems where quality is the priority.
  • xhigh is the new level — recommended for coding and agentic tasks. Claude Code has raised its default to xhigh across all plans.
  • max provides the deepest reasoning with no token constraint. It applies to the current session only (unless set via environment variable) and doesn't persist.
  • low​ and ​​**medium** trade accuracy for speed and cost. Useful for high-volume classification or routing where marginal quality differences don't justify spending.

One detail I went back and re-read twice: ​Opus 4.7 respects effort levels more strictly than Opus 4.6​, especially at low and medium. If you observe shallow reasoning on a complex task, the right move is to raise effort — not to add prompt scaffolding around it. The docs are explicit about this.

When running at xhigh or max, set ​max_tokens​​ to at least 64k to give the model room to think and act across subagents and tool calls. Starting at 64k and tuning from there is Anthropic's own recommendation.

Features That Matter for Long-Running Agents

High-Resolution Vision, xhigh Effort, and Tool Workflows

The vision upgrade is the most concrete capability gain for agent builders. As documented in Vellum AI's benchmark analysis of Opus 4.7, OSWorld-Verified (computer use) climbed from 72.7% on Opus 4.6 to 78.0% — a 5-point gain that, paired with the resolution upgrade, shifts the economics of UI automation meaningfully.

Opus 4.7 is ​the first Claude model with high-resolution image support​: maximum resolution increased from 1,568 pixels (~1.15MP) to 2,576 pixels (~3.75MP) on the long edge. That's more than triple the pixel budget. For computer-use agents that read dense UIs, screenshot-based workflows, or document understanding pipelines, this is a meaningful change. Critically, coordinates now map 1:1 with actual image pixels — the scale-factor math that was previously required for coordinate extraction is gone.

Vision in a request:

Python
message = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=4096,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {"type": "url", "url": "https://example.com/diagram.png"}
                },
                {"type": "text", "text": "List every service shown and the connections between them."}
            ]
        }
    ]
)

If the extra resolution isn't needed for a given task, downsample images before sending — high-resolution images produce more tokens, and for cost-sensitive workloads that adds up.

For tool-heavy agentic loops, raising effort increases the frequency and depth of tool calls. The relationship is direct: lower effort → fewer tool calls and shallower reasoning chains; higher effort → more thorough tool interaction. This is steerable through prompting as well, but the effort parameter is the cleaner lever.

Cost and Latency Controls for Production Use

Task budgets are the new mechanism for controlling spend on long agentic loops. Enable them with the beta header:

Python
response = client.beta.messages.create(
    model="claude-opus-4-7",
    max_tokens=128000,
    output_config={
        "effort": "high",
        "task_budget": {"type": "tokens", "total": 128000}
    },
    messages=[
        {"role": "user", "content": "Review the codebase and propose a refactor plan."}
    ],
    betas=["task-budgets-2026-03-13"]
)

The model sees the countdown and uses it to prioritize and wrap up gracefully. Without a task budget, default behavior is "spend as needed" — on a complex agentic task at xhigh, that can mean significantly more output tokens than you'd expect from a single-turn request.

For async workloads — evaluation runs, nightly summaries, batch analysis — the Batch API gives a 50% discount and removes rate-limit pressure from real-time traffic. Prompt caching remains available and can cut repeated input costs by up to 90% for workloads with stable system prompts or large static prefixes.

Migration Mistakes to Avoid

Assuming 4.6 Prompts Transfer Unchanged

This is the one I keep seeing come up.

Opus 4.7 follows instructions more literally than Opus 4.6. It no longer reads between the lines or silently generalizes from one case to another. Soft phrasings — "try to," "if possible," "roughly" — now carry more actual weight. Prompts that relied on 4.6's interpretive flexibility will sometimes behave differently, and not always in the way you'd expect.

The three breaking API changes require code updates, not just prompt edits:

  1. Replace thinking: {type: "enabled", budget_tokens: N} with thinking: {type: "adaptive"}
  2. Remove temperature, top_p, top_k from requests entirely
  3. Audit max_tokens — the new tokenizer maps the same text to up to ​1.35× more tokens​, so existing ceilings may cut off responses that previously fit

On tone: Opus 4.7 is more direct and opinionated than 4.6 — fewer emoji, less validation-forward phrasing. If your product relies on a specific voice calibrated to 4.6's warmer style, re-evaluate your style prompts against the new baseline before promoting to production.

Anthropic's Opus 4.7 launch announcement links directly to the full migration checklist, and Claude Code users can run /claude-api migrate to automate the model ID swap and breaking parameter changes across a codebase.

Treating Long Context Like Persistent Memory

I paused here when I read this in the docs, because it's easy to conflate the two.

A 1M token context window is not persistent memory. It means you can fit 1 million tokens into a single request. When that request ends — when the session closes, when the agent crashes, when a new conversation starts — that context is gone. The next request starts at zero.

Opus 4.7 does include improved file system-based memory: the model reads and writes to notes files across multi-session work, with noticeably more reliable behavior for agents using this pattern. But that's a tool you configure. It doesn't happen automatically from a long context window.

The distinction matters for anyone building agents that are supposed to "remember" things across runs. The 1M window helps within a session. Cross-session memory requires explicit architecture.

What the API Still Does Not Solve

Capability Reuse and Validated Repair History

This is the part I find harder to fit neatly into an API guide, but I think it's worth saying.

When Opus 4.7 catches a logical fault during the planning phase — which the release materials describe as a genuine capability — that reasoning happens within the session. The fact that it caught the fault, and how it corrected it, doesn't automatically persist as a reusable pattern for the next run. The next time the same class of problem appears, the model reasons from scratch.

According to a research paper on AI agent reliability published in early 2026, most models are benchmarked on average accuracy rather than consistency across runs — which means a model can score well on benchmarks while still failing unpredictably on the same class of task at different times. Opus 4.7 improves on Opus 4.6's baseline. That's real. But in-session correction and cross-session capability inheritance are different problems, and the API solves the first, not the second.

For builders running production agents, this means the engineering work of building reliable systems — evaluation harnesses, repair documentation, monitoring — still sits outside the model itself. The API gives you a more capable model. What you do with that capability over time is still on your architecture.

FAQ

Q: What is the correct model ID for Opus 4.7?

A: claude-opus-4-7. This is the stable model string for API calls across the Claude API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.

Q: Is adaptive thinking on by default?

A: No. Adaptive thinking is off by default on Opus 4.7. You must set thinking: {"type": "adaptive"} explicitly to enable it. Requests with no thinking field run without thinking.

Q: What happens if I send ​temperature​​​ or ​budget_tokens​?

A: Both return a 400 error on Opus 4.7. Remove temperature, top_p, and top_k from all requests. Replace budget_tokens with output_config: {"effort": "..."} and thinking: {"type": "adaptive"}.

Q: When should I use xhigh vs high effort?

A: Use xhigh as the starting point for coding and agentic tasks — it's now the default in Claude Code across all plans. Use high for most intelligence-sensitive tasks. Drop to medium or low when latency or cost is more important than reasoning depth. If you observe shallow output on a complex task at a lower level, raise effort rather than adding prompt scaffolding.

Q: Does the 1M context window mean the agent remembers things between sessions?

A: No. The context window applies within a single request. When a session ends, that context is gone. Multi-session memory requires explicit tooling — file-based memory, external stores, or similar architecture.

Previous Posts

Related Articles