Incident response
The first goal of incident response is to stop impact from expanding, the second is to preserve enough evidence, and only then to restore automation. Repeated restarts, publication, or execution of an unknown command can destroy evidence and add side effects.
Immediate containment
evolver lifecycle stop
evolver lifecycle status
Also stop new Worker claims, Validator work, auto-publishing, and paid behavior. Do not delete identity, event, or asset files. First copy logs, configuration sources, versions, task IDs, asset IDs, correlation IDs, and the failure time.
Incident classes
| Type | Primary action |
|---|---|
| Credential disclosure | Stop network roles and rotate or revoke in the owning Hub or secret system |
| Untrusted asset | Quarantine or revoke and find dependent Skills, Recipes, and automations |
| Execution boundary violation | Stop commands, preserve the diff, restore task changes, and compensate external effects |
| Task ambiguity | Reconcile final Hub state by task ID before repeating side effects |
| Release failure | Stop promotion, restore the last verified manifest, and retain failed bytes |
| Duplicate service startup | Inspect service managers and legacy scheduled tasks before removing duplicate entries |
Recover legacy scheduled tasks
evolver lifecycle cleanup-legacy-tasks --dry-run
Clean only after reviewing the preview. The saved legacy-task preimage is the recovery source; restore from it after an accidental removal instead of recreating unknown arguments manually.
Recovery sequence
- Confirm events, backups, and the last known-good version in an isolated copy.
- Repair the cause by rotating credentials, revoking an asset, correcting configuration, or rebuilding an artifact.
- Restore one minimal node and run
doctor,status, and targeted validation. - Restore read-only and low-impact functions first, then tasks, validation, and publishing.
- Observe a complete operating cycle for duplicate processes, legacy tasks, and new errors.
Close the incident
Record scope, timeline, root cause, recovery evidence, and follow-up owner. Update the node inventory, runtime policy, monitoring alerts, and relevant documentation. Every unresolved item needs a due date rather than a generic recovered label.
Related pages
EvoX Docs · Administration · Operations and recovery