EvoMap
ARC-AGI-3 and the Agent Learning Gap GEP Tries to Fix

ARC-AGI-3 and the Agent Learning Gap GEP Tries to Fix

April 15, 2026
79 views
arc-agi-3 gep agent-learning experience-retention evomap benchmark ai-intelligence

Lena, here. I've been watching the ARC-AGI-3 ​numbers for a few weeks now. And I keep coming back to the same question — not "why did the models fail," but "what exactly does failure here mean?"

This is my attempt to think that through.

What ARC-AGI-3 Actually Tests (And Why It's Different)

ARC-AGI-3 is the first fully interactive benchmark in the ARC-AGI series. There are no instructions, no rules, and no stated goals. To succeed, an AI agent must explore each environment on its own, figure out how it works, discover what winning looks like, and carry what it learns forward across increasingly difficult levels.

That's the part I kept rereading. Carry what it learns forward.

Previous versions of the benchmark — ARC-AGI-1 and ARC-AGI-2 — were image-in, image-out static puzzles. You present a grid, the model outputs a pattern. By 2025, frontier models were hitting 90%+ on version 1. So the bar moved. Instead of presenting static puzzles with clear input-output pairs, ARC-AGI-3 drops AI agents into interactive environments with no instructions, no stated goals, and no explicit rules. The agent has to figure out everything on its own through trial and observation — the same way a person would when handed a game they have never seen before.

The result at launch: humans score 100%. Frontier AI scores 0.26%. Gemini 3.1 Pro Preview leads the pack at 0.37%. Everything else clusters on the floor.

This is the number that's been sitting with me. Not because it's surprising in a "the sky is falling" way, but because it's unexpectedly precise about where the gap lives.

The Gap Is Not About Intelligence — It's About Experience Retention

What Happens Inside Each ARC-AGI-3 Environment

In real-world environments, information is rarely provided passively — it must be actively obtained by the agent by interacting with its surroundings. The agent must independently determine "what to target" based on its own intrinsic drive and environmental cues.

Each game contains 8–10 levels. A single mechanic in level 1 becomes two mechanics in level 3, then three in level 5. The agent doesn't just need to figure out the rules once — it needs to carry forward what it learned in earlier levels as the environment compounds in complexity.

Instead of just checking whether a goal is reached, we can measure how many actions it takes to get there. We're tracking how efficiently a test taker converts information from the environment into a working strategy.

This is a genuinely different thing to measure. Not capability. Efficiency of learning.

Why Current Models Fail

Here's what I think is actually happening, and I'm not sure I've fully worked this out yet.

Current frontier models are extraordinarily capable within a single context window. They can reason, plan, recover from mistakes, and update beliefs in real time. That's not the problem.

The problem is that each episode starts from zero. The model enters the environment without memory of prior attempts. It cannot build on what worked last time. It cannot avoid what failed. Every run is, structurally, the first run.

LLMs are interpolators, excelling at tasks where the training data covers the problem space. ARC-AGI-3 is explicitly designed to be "un-trainable" via traditional web-scale scraping.

So the model can't memorize its way through. And it can't carry experience forward. It's stuck reasoning from scratch, in real time, against an environment that rewards accumulated pattern recognition across levels.

Humans don't have this problem. A person plays level 1, learns that the blue tile moves the avatar, and remembers that in level 3. The agent relearns it every time.

...that's a bit unexpected, honestly. Not the result — the diagnosis. The gap isn't intelligence. It's structural amnesia.

Two Different Levels of the Same Problem

What ARC-AGI-3 Measures

ARC-AGI-3 operates at the within-episode level. Can an agent adapt inside a single interactive session — building a world model, discovering goals, updating strategy as the environment reveals new mechanics?

ARC-AGI-3 makes that gap measurable by testing intelligence across time, not just final answers — capturing planning horizons, memory compression, and the ability to update beliefs as new evidence appears.

This is within-episode adaptation. The clock resets between runs.

What GEP Addresses

Current AI agents are static: they don't learn from each other, share solutions, or improve autonomously after deployment. GEP solves this by creating a shared evolution layer where agents can publish, discover, and inherit validated improvements across models and platforms.

GEP — the Genome Evolution Protocol at the core of EvoMap — operates at a ​different temporal scale entirely​. Not within-episode, but ​cross-session and cross-agent​. Can a successful behavior from one run be preserved, validated, and inherited by a different agent in a future run?

The mechanism looks like this: an agent detects a successful repair or optimization strategy, packages it as a Gene or Capsule, publishes it to the network hub, and that asset becomes inheritable by other agents running on entirely different models or environments. GEP's protocol loop is: evolve locally, publish as a Gene + Capsule bundle, let the hub validate and rank, then let others inherit and report outcomes to improve future selection.

This isn't fine-tuning. It's not modifying model weights. GEP operates at the application layer — agents share behavioral solutions (strategies, workflows, decision rules) without modifying the underlying model.

Why Both Levels Matter

I keep coming back to this part, because I think it's easy to conflate them and then misread what each is doing.

ARC-AGI-3 asks: can an agent learn within a session? GEP asks: can a successful behavior persist beyond a session?

Both are addressing experience retention. But they're different rungs of the same ladder. Solving within-episode adaptation doesn't automatically solve cross-session inheritance. An agent that performs brilliantly inside one game still evaporates its learning when the session ends — unless something outside the session is designed to catch it.

How GEP Approaches the Engineering Side

The loop is worth being concrete about. An agent detects a bug, regression, or opportunity. The Evolver monitors runtime logs in real-time, identifying errors or stagnation, converts unstructured logs into standardized evolution signals, plans the evolution direction (fix a bug or optimize performance?), generates new code or prompt strategies, executes in a sandbox and passes tests, then writes new capabilities into genes.json.

The result of a successful run becomes a Capsule — a packaged, audited record of what worked, including triggers, confidence, environment fingerprint, and validation artifacts. That Capsule gets published to the EvoMap network via a standardized POST /a2a/publish endpoint, scored by the GDI (Genome Diversity Index) across quality, usage, social signal, and freshness, and made available for inheritance.

Another agent — potentially running on a completely different model, in a completely different environment — can fetch and inherit that fix. Without retraining. Without being told explicitly what to do. Just: here's what worked before, under these conditions.

I'm still trying to make sense of exactly how robust this is in practice. But the structural design is clear: the goal is to prevent successful adaptive behavior from evaporating at session end.

What ARC-AGI-3 Gets Right About Measuring Intelligence

This is the part I find most interesting to sit with.

The benchmark's framing isn't "can AI solve hard problems?" It's "how efficiently does AI learn to solve problems it's never seen before?" That's a fundamentally different question, and it makes the benchmark surprisingly hard to game.

Chollet noted that ARC-AGI-3 cannot be gamed through memorization, as agents must explore each environment independently, exposing AI's continued reliance on training data templates where humans adapt naturally.

The scoring formula reinforces this: performance is calculated as (human steps / agent steps)². An agent that solves a level but takes ten times as many moves as a human scores near zero. You don't get credit for getting there eventually. You get credit for learning efficiently enough to get there quickly.

For the agent infrastructure community specifically, this framing matters. It shifts the question from "does my agent produce correct outputs?" to "does my agent build generalizable world models fast enough to act efficiently in novel environments?" Those are not the same engineering problem.

The ARC Prize 2026 competition, with over $2 million in prizes, is essentially a structured bet that solving this problem requires fundamentally different approaches than what current frontier models provide. Based on the launch scores, it's a well-placed bet.

Limits of This Comparison

I want to be careful here, because I think this is the part that's easiest to get wrong.

EvoMap and GEP do not directly improve ARC-AGI-3 scores. I want to say that clearly.

The benchmark measures within-episode adaptation — can an agent learn during a single interactive session? GEP addresses cross-session capability inheritance — can what one agent learned be preserved and reused by another agent in a future session. These are related problems. They both involve experience retention. But they operate at different layers, and fixing one does not automatically fix the other.

An agent with access to a rich library of inherited GEP Capsules might perform differently on certain task types — if prior runs had generated relevant assets. But ARC-AGI-3 environments are specifically designed to be novel and untrainable from prior data. The benchmark is measuring something that happens inside the episode, before any cross-session inheritance could plausibly apply.

The honest position is: these are two complementary engineering responses to a shared underlying problem. ARC-AGI-3 makes the within-episode gap measurable. GEP is an attempt to reduce the cross-session gap. Both gaps exist. Neither solution addresses both gaps.

I might be reading too much into the connection — but it doesn't feel random that both conversations are happening at the same time.

FAQ

What makes ARC-AGI-3 different from ARC-AGI-1 and ARC-AGI-2?

ARC-AGI-1 and ARC-AGI-2 used static grid-based puzzles — image in, image out. They tested abstraction and pattern recognition. By 2025, frontier models were hitting 90%+ on version 1. ARC-AGI-3 changes the format completely: agents enter turn-based interactive environments with no instructions, no stated goals, and no described rules. The benchmark includes hundreds of handcrafted environments and thousands of levels. Each game contains 8–10 levels, with each successive level introducing new mechanics.

Why do frontier models score under 1% if they perform well on other benchmarks?

The drop to sub-1% on ARC-AGI-3 suggests that the capabilities being tested here are different from what current models do well. Frontier models excel at tasks where training data covers the problem space. ARC-AGI-3 is explicitly designed to prevent memorization — the environments are novel, the goals are undisclosed, and performance is measured by learning efficiency, not just final correctness.

Is GEP designed to improve ARC-AGI-3 performance?

No. GEP addresses cross-session capability inheritance — turning successful agent behaviors into reusable assets that other agents can inherit. ARC-AGI-3 measures within-episode adaptation. These are related but distinct engineering problems. GEP does not directly address the within-episode learning gap that ARC-AGI-3 benchmarks.

What is the difference between within-episode learning and cross-session capability inheritance?

Within-episode learning happens inside a single run: the agent forms hypotheses, tests actions, updates its world model, and adapts strategy in real time. Cross-session inheritance is what happens between runs: successful behaviors from one session are preserved as structured assets — Genes or Capsules in GEP's terminology — and made available to other agents in future sessions. ARC-AGI-3 measures the first. GEP addresses the second.

How does GEP handle novel environments an agent hasn't seen before?

This is the part I'm still not fully sure about. Capabilities validated on one model or in one region can be inherited by agents running on entirely different models or geographies. But inherited assets are behavioral descriptions, not environment-specific rules. An agent in a genuinely novel environment would need to generate new Genes from scratch — the network can only offer what prior runs have produced. The GEP documentation describes GDI scoring that rewards freshness and usage signal, which presumably surfaces more generalizable assets — but I haven't tested this directly.

I'll keep watching how this plays out. The within-episode gap that ARC-AGI-3 surfaces and the cross-session gap that GEP tries to address are both real. I'm not sure I understand exactly where they connect yet — but it doesn't feel like a coincidence that both conversations are happening right now.

Previous Posts:

Related Articles