Protocol · formal layer
GEP: what turns a prompt fragment into an evolvable object
The Gene Evolution Protocol is the layer of arXiv:2604.15097 that makes the strategy gene more than a well-written snippet. Left un-protocolized, a gene is free-form text. Its boundaries are unstable, its fields hard to compare, its revision history indistinguishable from noise. Canonicalized under GEP, it becomes an explicit object — one that can be matched, replaced, revised, validated, and audited like any other engineered artifact.
The paper introduces GEP in Appendix A and keeps the experiments focused on the gene layer, but the protocol defines three object types, and the division of labor is the point:
Object hierarchyThree layers, three jobs
| Object | Role | What it holds | Answers the question |
|---|---|---|---|
| Gene | Control unit | Matching signals, summary, strategy steps, AVOID cues; optional constraints and validation hooks | What strategy should govern this kind of task? |
| Capsule | Validated execution unit | Task signature, genes instantiated, execution trace, outcome, validation record, lineage pointer | How was a strategy actually realized and validated? |
| Event | Evolution record | Event type (repair, innovation, validation_pass, validation_fail, solidify), source and target assets, trigger, diff, timestamp | How did this capability come to be, and who changed it? |
Structure as formalized in Appendix A of arXiv:2604.15097. Capsules capture successful compositions; events are immutable and exist for provenance, not for control.
The separation matters because each layer has a different failure cost. A bad gene is a bad prompt — cheap to retire, provided events record why it was retired. A bad capsule is a false precedent. An event log that can be edited is no audit at all. Hence the asymmetry: genes are designed to change, events are designed never to.
The loopSix stages from runtime noise to solid asset
GEP updates experience through a six-stage loop the paper calls the protocol-level analogue of trial–validation–solidification. Its function, in the paper’s words, is not to repair a single run but to convert transient runtime adaptation into persistent reusable experience.
- 1 · ScanWatch runtime traces, tool logs, execution failures, stagnation signals.
- 2 · SignalConvert raw traces into standardized protocol signals — machine-actionable mutation and repair triggers.
- 3 · IntentDecide the objective: repair, optimization, or extension.
- 4 · MutateGenerate a candidate asset — rewrite strategy steps, AVOID cues, structure, or metadata.
- 5 · ValidateExecute the candidate in a sandbox or check it against validation hooks. Only validated assets may persist.
- 6 · SolidifyWrite the validated capability back as a new or revised gene; update capsules and event records.
VersioningA gene is born, tested, promoted — or rolled back
The loop’s payoff is lifecycle discipline. Because every revision is a discrete, validated object, a capability has a version history rather than a blob: mutations are proposed, validation decides, and solidification promotes. The event types the protocol logs — mutation, validation_pass, validation_fail, solidify — are exactly the states a version travels through.
The protocol also lives outside the paper. Evolver — the evolution engine the paper pairs with OpenClaw for its CritPt runs — is maintained by EvoMap, and their essay comparing agent skills with GEP genes argues the engineering-side contrast from a different angle than this guide: developer-registered tools such as Semantic Kernel plugins or LangChain actions are static once shipped, whereas a gene carries its validation history, mutates when it fails, and chains into capsules. Worth reading alongside Appendix A; note that the “skill” in their comparison is the tool/plugin sense, not the documentation-style Skill packages the paper benchmarks.
InvariantsWhat the protocol refuses to admit
Appendix A.7 lists the invariants a protocolized object must satisfy, and they read like a code-review checklist for experience:
- Stable boundaries — explicit fields, canonical serialization; an object, not a snippet.
- Control-oriented structure — compact control content, not documentation-heavy explanation.
- Operability — matchable, replaceable, revisable, composable.
- Validatability — executable or explicit validation interfaces.
- Lineage and auditability — enough provenance to reconstruct how an object was produced and changed.
In the evolutionary runs reported on CritPt, this discipline is what shows up in practice: repair genes that package a constrained loop — diagnosis, blast-radius estimate, smallest reversible patch, validation, solidification — and evolution events that log each attempt with its outcome, failures included. The findings page carries the numbers: paired base models lifted from 9.1% to 18.57% and from 17.7% to 27.14% across the two runs, with weights untouched.
One more detail from the paper’s appendix is easy to skim past and worth ten seconds: in some analyses the gene is wrapped in an evolution-style context that presents its validated history to the model — four failed attempts, then the strategy that passed. The history is not replayed as a log; it is summarized as evidence that the strategy earned its place. That framing is the protocol speaking to the model.