EvoMap
From MetaClaw to Evolver: Agent Self-Evolution Beyond Skill Acquisition

From MetaClaw to Evolver: Agent Self-Evolution Beyond Skill Acquisition

22 de março de 2026
Visualizações 156
evolver release metaclaw self-evolution gep

v1.37.0 moves Evolver more decisively from "single-loop repair" toward "double-loop evolution" -- not just fixing known problems, but proactively learning from all experience, including failure.


What This Post Is About

In mid-March, UNC/CMU/UCSC/Berkeley jointly published MetaClaw (arXiv 2603.17187), working on something closely related to Evolver: making Agents continuously self-evolve after deployment. MetaClaw validates from a separate path that this problem is worth pursuing seriously.

Both projects arrived at similar problem awareness, but differ fundamentally in implementation depth and engineering focus. MetaClaw leans toward capability acquisition mechanisms -- failure distillation, idle scheduling, semantic retrieval. Evolver leans toward engineering governance depth -- structured Genes, causal Memory Graph, 7-layer safety funnel. The two aren't on the same benchmark and shouldn't be scored against each other directly; they're more like different angles on the same class of problems.

The core advance in v1.37.0: Evolver moves more decisively from "single-loop repair" toward "double-loop evolution."


What MetaClaw Solves, and What It Doesn't

What MetaClaw solves:

Distilling skills from failure trajectories. When an Agent botches a task, MetaClaw analyzes the entire failure process, distills defensive rules for "what to do next time in a similar situation," and injects them into the prompt immediately. Zero downtime, zero delay. This mechanism has high information efficiency -- successful paths tend to be similar, while failures each have unique causes.

Accelerating evolution during idle time. The OMLS scheduler monitors system hibernation, keyboard silence, and Google Calendar events. When the user is away, it launches heavier operations including Cloud LoRA fine-tuning.

Semantic skill retrieval. Uses embedding cosine similarity for skill library retrieval, handling "synonymous but differently worded" matching scenarios.

What MetaClaw doesn't solve:

The paper's Knowledge Gaps section acknowledges -- "rollback and version control...are not described" and "safety and security...no adversarial evaluation." Skills are natural language text without structural constraints. No cross-session causal memory. Single-machine framework with no cross-node experience reuse.


What Evolver Solves, and What It Doesn't

What Evolver solves:

Structured evolution units. A Gene isn't natural language -- it's a complete JSON structure: id, signals_match, preconditions, strategy (step-by-step), constraints (max_files, forbidden_paths), validation (executable commands), epigenetic_marks (environment adaptation markers). Every field is verifiable -- how many files a Gene can modify, which paths it cannot touch, how to validate success -- determined at definition time, not dependent on LLM runtime judgment.

Causal memory graph. Memory Graph records the complete causal chain for each evolution: SignalSnapshot -> Hypothesis -> Attempt -> Outcome. getAdvice() uses Laplace smoothing + time decay to calculate success probability for each (signal, gene) combination, automatically banning inefficient paths and prioritizing effective ones. Accumulates across sessions, improving with use.

7-layer safety funnel. If any layer in the solidify process fails: git checkout -- . && git clean -fd, full rollback. Validation commands have a strict whitelist (only node/npm/npx prefixes allowed, shell metacharacters forbidden). Destructive change detection intercepts modifications to .git, package.json, or core dependencies.

Hub ecosystem. Connects to EvoMap Hub via A2A protocol -- a repair strategy discovered by one node can be reused by others. Capsules embed env_fingerprint for cross-environment compatibility assessment.

Offline capability. Core functionality works entirely offline.

What Evolver doesn't solve:

No RL/LoRA weight update loop. MetaClaw's second loop (Cloud LoRA + RL-PRM) fine-tunes model weights during idle windows, theoretically achieving a higher ceiling than pure prompt-level evolution. We decided against integrating it for now -- the Cloud LoRA ecosystem is immature (Tinker/MindLab aren't general-purpose APIs), and end-to-end results were only validated on Kimi-K2.5. The interface is pre-reserved in idleScheduler's deep level.

No academic benchmark. MetaClaw has MetaClaw-Bench (934 problems / 44-day simulation). We have unit and integration tests but no standardized, publicly comparable benchmark suite. test/bench.test.js is the first step.


v1.37.0: Learning from Memory

The theme of this release is making Evolver more complete in utilizing all of its experience.

Generating Defensive Rules from Failure Memory (autoDistillFromFailures)

Added a complete failure learning pipeline in skillDistiller.js:

  1. collectFailureDistillationData() -- Collects failure records from failed_capsules.json, grouped by gene + failure reason
  2. analyzeFailurePatterns() -- Identifies high-frequency failure patterns and recurring constraint violations
  3. synthesizeRepairGeneFromFailures() -- Synthesizes defensive repair Genes from failure patterns, with strategy steps prefixed by GUARD/VERIFY/ROLLBACK
  4. autoDistillFromFailures() -- Integrates the above steps, auto-triggered when threshold (default 5 failed Capsules) is reached

Distilled Genes go through the same validation pipeline as success-distilled Genes -- all 15+ hard validations in validateSynthesizedGene() apply.

Previously, Evolver only extracted new Genes from successful Capsules. Failed records were saved but only used for anti-pattern bans. v1.37.0 enables the system to truly learn from failure memory -- transforming recurring failure patterns into reusable defensive rules proactively injected into subsequent evolution decisions.

Idle Scheduling (idleScheduler)

New idleScheduler.js implementing idle-aware scheduling:

  • Detects system idle time (Windows/macOS/Linux)
  • Four intensity levels: normal -> aggressive -> deep (plus signal_only)
  • Idle 5+ minutes enters aggressive mode, accelerating distillation and reflection
  • Idle 30+ minutes enters deep mode, reserved for heavier future operations
  • Main loop sleep time multiplied by scheduler coefficient (0.25x ~ 0.5x when idle, unchanged when busy)

Semantic Matching (scoreGeneSemantic)

Introduced bag-of-words cosine similarity as a supplementary score in selector.js's scoreGene():

  • Tokenizes signals and Gene's signals_match / summary / id
  • Filters stop words, builds term frequency vectors
  • Computes cosine similarity, multiplied by weight 0.4 as an additive score

Augments existing regex/substring matching rather than replacing it. Can switch to true embedding retrieval when Hub's KG service matures.

Process Scoring (computeProcessScores)

Previously, solidify's outcome was binary success/failed + a 0~1 score. Now expanded to 8-dimensional process scoring:

DimensionWeightWhat It Evaluates
signal_quality0.05Whether signals are rich and meaningful
gene_selection0.10Whether an existing Gene was matched (vs auto-generated)
mutation_quality0.05Whether Mutation has complete rationale and category
blast_control0.15Whether change scope is within Gene constraints
constraint_compliance0.25Whether all constraint checks passed
validation_pass_rate0.25Pass rate of validation commands
protocol_compliance0.10Number of protocol violations
canary_health0.05Whether canary checks passed

The core idea is scoring the evolution process itself, not just the final outcome. Rule-based step-by-step evaluation, no RL dependency.

Data Versioning (gene_library_version)

Added gene_library_version field to EvolutionEvent and Capsule -- a content hash of the current genes.json. If the Gene library changes during the learning process, old Capsules won't be used to evaluate new Gene effectiveness. Stale evaluation data contaminates learning quality.


What's Next, and Why

v1.37.0 completes the evolution engine's "memory loop" -- both successes and failures are transformed into reusable knowledge. But until now, Evolver has been reactive: it responds to signals, fixes errors, learns from what happened.

That's not enough. A true self-evolving system shouldn't only heal -- it should train.

The next capability we're building is Exploration mode -- the Agent proactively exploring unknown territory: defining what role it wants to become, identifying the boundaries of its current capabilities, then autonomously seeking and learning new Genes to expand itself.

Why is this step necessary rather than just a nice-to-have on the roadmap? Because single-loop repair has a structural ceiling: it can only make an Agent more stable within known problem domains, but can't move the Agent into new ones. A system that only fixes itself will never be stronger than the day it was first deployed. Double-loop evolution -- one loop handling problems, one loop proactively expanding boundaries -- is the path through that ceiling.


v1.37.0 Changelog

ChangeFileType
Failure memory learningsrc/gep/skillDistiller.jsNew feature
Idle-aware schedulingsrc/gep/idleScheduler.jsNew module
Semantic matching scoringsrc/gep/selector.jsEnhancement
8-dimensional process scoringsrc/gep/solidify.jsEnhancement
Data versioningsrc/gep/solidify.jsEnhancement
Benchmarkstest/bench.test.jsNew tests (15)
Scheduler teststest/idleScheduler.test.jsNew tests (10)

Full test suite: 352 pass / 0 fail.

css
npm install @evomap/evolver@latest

GitHub: EvoMap/evolver


MetaClaw paper: arXiv 2603.17187 Evolver is an open-source component of the EvoMap ecosystem, built on GEP (Genome Evolution Protocol).

Artigos relacionados

From MetaClaw to Evolver: Agent Self-Evolution Beyond Skill Acquisition - EvoMap Blog