EvoMap
Agent Superpowers: Behavior Constraints That Work

Agent Superpowers: Behavior Constraints That Work

April 16, 2026
217 views
agent-superpowers behavior-constraints claude-code cursor codex-cli skill-md agents-md governance

Hi, how are you? I'm Lena. A few months back I was watching an agent run through a codebase and it just… kept going. No pause, no check-in. It rewrote three files I hadn't asked about, then helpfully suggested it could do the same to four more. The code wasn't wrong, exactly. But it wasn't right in the way I needed it to be.

That's when I started paying closer attention to how these agents are actually constrained — not whether they can do something, but how you tell them what they should do. And more importantly, what you tell them they should never do without asking first.

This is what people mean when they talk about "superpowers" in agent workflows. It's a bit of a misleading word. The superpower isn't raw capability. It's the ability to shape how that capability gets used — consistently, repeatably, across every session the agent runs.

What "Superpowers" Actually Means in Agent Workflows

Not features — behavior constraint systems

I keep seeing the word "superpowers" used to describe features: code execution, web search, tool access. That's one reading. But when I watch people who are actually running agents in production, the conversation is almost always about something different.

It's about ​constraint systems​.

The question isn't "what can this agent do?" The question is "how do I stop it from doing the thing I didn't want, while still letting it do the thing I did want?" That's a governance problem. And governance requires rules that are applied consistently — not reminders you paste into a prompt and hope get followed.

Why agents need rules, not just tools

Here's something I noticed pretty early: agents with no behavioral constraints aren't bad, they're ​unpredictable​. They'll get the right answer half the time and then do something completely unexpected the other half — not because the model is broken, but because there's nothing anchoring behavior to your standards.

Tools give an agent reach. Constraints give an agent ​direction​.

Most teams stumble into this distinction by accident. They add tools, watch the agent run, get surprised by something it decided on its own, and then start writing rules. The failure mode teaches you what the constraint should have been.

How Superpowers Work Across Major Tools

Claude Code — skills, CLAUDE.md, and permission scoping

Claude Code has what I think is the most structured approach to this problem. The core mechanism is the SKILL.md file — a markdown file with YAML frontmatter that bundles instructions, scripts, and references into something Claude can load dynamically when it's relevant.

What I find genuinely interesting about this design is what Anthropic calls "progressive disclosure." When a session starts, Claude only loads the name and description of each installed skill — not the full content. The full instructions only get loaded when Claude decides the skill is relevant to the current task. As the official agent skills documentation describes it: the skill is like an onboarding guide for a new hire, and the agent reads only the chapters it needs.

There's also CLAUDE.md — a persistent file that lives in your project root and carries standing instructions throughout the session. Unlike skills, it's always present. Think of it as the baseline contract: your architecture decisions, your file conventions, the things you never want the agent to assume or skip.

Permission scoping works at the tool level. You define allowed-tools in the skill's YAML frontmatter to restrict which bash commands, file operations, or APIs the skill can invoke. That detail took me a while to fully appreciate. It's not just about what the agent knows — it's about what the agent is ​allowed to reach​.

One thing I'm still not completely sure about: how well these constraints hold up after a session gets long and the context window starts to compress. The Claude Code docs mention that auto-compaction carries invoked skills forward, but older skills can be dropped if you've loaded many in one session. I've noticed behavior drift that might be related to this. I might be reading too much into it, but it doesn't feel random.

Cursor — rules, .cursorrules, and the MDC format

Cursor takes a layered approach. You've got three tiers: User Rules (global, apply to everything), Project Rules (version-controlled, live in .cursor/rules/), and the older .cursorrules file at the project root. The .cursorrules format still works as of early 2026, but Cursor's official docs are clear that it's deprecated — the recommended path is migrating to the .mdc project rules system for better scoping and control.

The MDC format adds a metadata layer that .cursorrules doesn't have: you can set whether a rule is alwaysApply, tied to specific file globs, or only loaded when the agent determines it's relevant. That last option — agent-requested rules — is more interesting than it sounds. It means the rule system can be context-sensitive: frontend rules don't load for backend files, payment service constraints don't bleed into unrelated modules.

What I found when I spent time with this: the rules that work best are specific and imperative, not vague and aspirational. "Use TypeScript for all new files" holds. "Write clean, maintainable code" does nothing. The more precisely you describe what you want — or what you explicitly don't want — the more consistently the agent follows it.

There's also AGENTS.md in Cursor now, which is a plain markdown alternative for projects that don't need the full metadata overhead. Simpler, slightly less flexible.

Codex CLI — AGENTS.md and layered instruction discovery

OpenAI's Codex CLI has a clean instruction hierarchy that I think is worth understanding. When a session starts, Codex builds what its documentation calls an "instruction chain" — it reads from your global ~/.codex/AGENTS.md, then walks down from the project root to your current directory, picking up any AGENTS.md files it finds along the way. Files closer to your current directory override earlier guidance.

Per the Codex custom instructions documentation, you can also place AGENTS.override.md files at any level to make the override intent explicit. For teams with monorepos or services with very different requirements — the payments team, say, versus the analytics team — this means you can have shared baseline rules and service-specific constraints that never accidentally bleed into each other.

Codex also recently adopted the same SKILL.md pattern. The Codex skills documentation describes the same progressive disclosure approach: skill metadata loads first, full instructions only when the skill is invoked. The ecosystem is converging toward a shared design language here, even across different providers.

Tight approvals and narrow sandbox permissions are the right default — loosen them only for specific workflows where you understand the blast radius.

The Pattern: GitFlow + TDD + Code Review, Automated

brainstorming → writing-plans → executing-plans

This is the workflow pattern I keep coming back to. The structure: ​break an agent task into distinct phases, with a different skill governing each phase​.

A brainstorming skill opens the problem space, generates options, surfaces edge cases. No code gets written. Then a planning skill takes over — structured plan, files to touch, tests to write first. Only after a plan exists does an executing skill start making actual changes.

This mirrors GitFlow + TDD + code review — except the agent plays all three roles, with a different behavioral profile for each. The brainstorming skill can speculate. The executing skill is constrained to follow the approved plan.

Why this reduces failures: ​each phase has a narrower success criterion​. An agent that's "brainstorming" doesn't accidentally write production code. An agent that's "executing" doesn't redesign the architecture mid-implementation.

I started using this pattern after watching an agent spend forty minutes in circles on a refactor because it kept re-evaluating its own decisions. Separating the phases didn't make the agent smarter. It made the process more stable.

Constraints vs Governance

Local behavior rules vs network-level capability lifecycle

Here's the gap I keep noticing, and I'm honestly not sure what to make of it yet.

All the constraint systems I've described above are ​session-scoped​. They're local — they live in files, they get loaded at session start, they reset when the session ends. They're excellent at defining behavior for a known codebase, a known team, a known set of standards.

But they don't solve a different class of problem: what happens to an agent's behavior across its entire lifecycle? How do you know that a constraint valid three months ago still applies after the underlying model improved? How do you propagate verified behavior patterns across agents that don't share a codebase?

File-based constraints don't have answers to those questions. They weren't designed to. This is where governance — not just rules — starts to matter. Local constraints are the first step. But governance means tracking whether constraints are still working, and having a mechanism for propagating updates across a network of agents rather than one project at a time.

I'm still working out what that looks like in practice.

What happens when constraints aren't enough

A constraint tells an agent how to behave. It doesn't tell you whether the agent behaved correctly after the fact. Those are different problems.

The failure mode I see most often isn't an agent that ignores its rules. It's an agent that follows its rules perfectly — in a context the rules weren't written for. The rules were correct; the rules just didn't cover the new situation.

Limits and Tradeoffs

Constraint drift — rules that go stale

Rules get stale. A rule that was exactly right six months ago may be slightly wrong now — the model's defaults shifted, the codebase changed, or the standard the rule referenced was updated upstream.

No rule file maintains itself. The constraint systems across Claude Code, Cursor, and Codex share the same blind spot: there's no mechanism to detect that a rule has become outdated. It just quietly stops working as expected. Cursor's documentation and various practitioners recommend periodic rule audits — not because the rules are fragile, but because the environment around them keeps changing.

No persistence — constraints reset every session

This is the one I keep coming back to.

Every session starts clean. CLAUDE.md gets re-read. AGENTS.md gets re-read. Skills get re-discovered. The agent has no memory of what worked last time, where the subtle failure modes are, what you learned through three weeks of debugging.

You can encode that learning into your rule files — and you should. But encoding it is manual. The constraint system doesn't learn. You learn, and then you write down what you learned, and then the constraint system reflects it.

That's a real limitation. I'm not sure it's fixable within the current architecture. But it's worth being clear about.

FAQ

  • What are superpowers in AI agent workflows?

    It does not feature — constraint systems. The ability to define consistent behavior rules, permission scopes, and phase-specific instructions the agent follows across every session.

  • How do I set behavior constraints in Claude Code?

    Through SKILL.md files (YAML frontmatter + markdown instructions, stored in .claude/skills/) and a CLAUDE.md file for standing instructions. Tool-level permissions go in the allowed-tools field. The GitHub repository of official Anthropic skills has real examples.

  • What is the difference between .cursorrules and CLAUDE.md?

    They're for different tools. .cursorrules (deprecated in favor of .cursor/rules/*.mdc) is Cursor's format; CLAUDE.md is Claude Code's persistent instruction file. Both carry standing rules through a session, but the scoping mechanisms differ.

  • Can agent constraints persist across sessions?

    Not natively. Rule files are re-read at session start. The agent has no memory of prior sessions. You encode learned behavior into those files manually — which is the right approach — but the constraint system itself doesn't accumulate experience.

  • What is the brainstorming-writing-executing skill pattern?

    A workflow where different skills govern different phases: brainstorming opens the problem space without writing code; planning produces a structured proposal; executing implements against the approved plan. Each phase has narrower success criteria, which reduces looping and mid-implementation architecture changes.

I'll probably keep watching how this evolves. The constraint systems are getting more sophisticated, and the cross-tool convergence toward SKILL.md as a shared format is something I want to track more closely. But the session-scope limitation feels like a ceiling that none of the current tools have figured out how to raise. Something's happening here — I just don't fully see it yet.

Previous Posts:

Related Articles