Concepts · definitions first
Skill vs Gene: a documentation artifact and a control object
A procedural Skill is a documentation-oriented experience representation: a human-readable package that organizes prior problem-solving knowledge into overview, workflow, API notes, examples, error handling, and pitfalls. It is built to be read, taught, reviewed, and archived.
A Strategy Gene is a control-oriented experience representation: a compact object distilled from that same experience, carrying task-matching keywords, a one-sentence summary, a short strategy list, and failure-aware AVOID cues, built to change what the model does on the next run under a tight token budget.
Those two sentences carry the whole page. In arXiv:2604.15097, the same underlying experience is cast into both forms and compared across 4,590 controlled trials: the gene ends at a 54.0% average pass rate, the full Skill at 49.9%, no guidance at 51.0%. The argument of this page is that the gap is neither noise nor brevity; it follows from what each encoding is for.
Side by sideEvery axis on which they differ
| Axis | Procedural Skill | Strategy Gene |
|---|---|---|
| Orientation | Documentation-oriented — written for human reading | Control-oriented — written for model-facing inference |
| Goal | Documentary completeness; teach and archive the process | Signal density under a constrained token budget |
| Organization | Documentation logic: overview → workflow → reference material | Control logic: when it applies → what to do → what to avoid → how to check |
| Failure knowledge | Recorded as history — error logs, pitfalls sections, worked failures | Compressed into explicit AVOID cues attached to the strategy |
| Structure’s role | Formatting serves readability | Editable schema is part of the effect — flattened to prose, the gain collapses (54.0% → 50.5%) |
| Accumulation | Grows by appending; history dilutes control (Skill + failure: 47.8%) | Evolves by selective revision; validated warnings attach cleanly (Gene + failure: 52.0%) |
| Natural consumer | A developer six months later | A model in the next run |
Percentages from the corresponding tables of arXiv:2604.15097, annotated on the findings page.
AnatomyWhat is actually inside each one
Concretely, for the paper’s running example (scenario S012_uv_spectroscopy — detect and measure peaks in UV-Vis spectra): the Skill side is the full package a maintainer would recognize, seven sections deep. The gene is eight lines. Both encode the same hard-won lesson about unit conversion; only one of them surfaces it as an instruction the model cannot miss.
The recurring failure the gene guards against is easy to underestimate. scipy.signal.find_peaks expects min_distance in sample-index units; a model that passes a wavelength value through unconverted produces plausible-looking output with wrong peak counts. Separately, reporting FWHM requires converting peak_widths output back to wavelength units first. Neither error breaks the program loudly. Both quietly cost checkpoints — which is exactly the kind of failure a compact AVOID cue is for.
SKILL.md · excerpt, ≈2,500 tokens in full
strategy-gene · complete, ≈230 tokens
The gene schema, field by field
Under the Gene Evolution Protocol, the gene is serialized as a structured object. The paper (Appendix A.3) lists the fields: type, schema_version, id, signals_match — the keywords that decide when the gene applies — then summary, strategy, optional constraints, optional validation hooks, and an asset_id for lineage. Each field earns its place in the controlled trials: keywords alone carry +2.5, but the jump to the strongest average comes when the strategy layer completes the object.
What the schema buys is operability. Genes with stable boundaries can be matched against a task, replaced by a revised version, composed deliberately rather than heaped, and validated before solidification. A prose blob can be none of those things — it can only be re-read, or re-generated. That is the precise sense in which a gene is not a shortened skill. It is the smallest unit of experience a system can operate on rather than merely quote.
Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package. The anatomy above is the paper’s explanation of why the small object wins.
For the protocol layer that manages genes over time — capsules, events, the six-stage loop — continue to the GEP page. For every table behind the percentages on this page, see findings.