Benchmark · the setting
Gene-Bench: where the 4,590 trials happen
Gene-Bench is the name the paper gives its benchmark setting: 45 scientific code-solving scenarios on which every condition in arXiv:2604.15097 is evaluated, for 4,590 retained trials in total. Each scenario asks the model to write a Python program; the program runs in a sandbox; a scenario-specific test script scores it against a predefined set of checkpoints. Across 4,590 controlled trials on 45 scientific code-solving scenarios, the compact Strategy Gene representation reached a 54.0% average pass rate, against 51.0% with no guidance and 49.9% with the full documentation-style Skill package.
The domain spread is deliberately wide: bioinformatics, neuroscience, chemistry, seismology, climate and atmospheric science, signal processing, network analysis, finance, robotics, and quantum computing. Task shapes vary just as much: structured parsing and transformation, analysis pipelines, simulation and fitting, algorithmic planning, structured output generation. That breadth is what licenses the paper’s representational claim beyond any single niche.
Worked examplesFour scenarios the paper puts on the table
The paper details four scenarios enough to picture the whole bench. They are worth quoting because they show why checkpoint scoring is not a technicality: these tasks have many independent ways to be partially right.
| Scenario | Domain | Checkpoints | What the program must do |
|---|---|---|---|
| S012_uv_spectroscopy | Chemistry / signal processing | — | Read UV-Vis spectra from CSVs, detect peaks, compute wavelength, height, FWHM and area, identify dominant peaks, write structured output. The paper’s running example. |
| S002_spike_behavior | Neuroscience | 12 | Read spike times and velocity from a MATLAB .mat file, filter successful trials, bin spikes, align behavior by interpolation, write trial-structured HDF5 with quality-control flags. |
| S026_earthquake_catalog | Seismology | 14 | Aftershock identification, completeness magnitude, Gutenberg–Richter b-value via the Aki maximum-likelihood formula, structured CSV/JSON summaries. |
| S114_obstacle_avoidance | Robotics | 11 | Build a valid 2D path via RRT or potential fields, geometric collision checking, path smoothing, and metrics: length, waypoints, clearance, smoothness. |
Checkpoint counts as stated in the paper; S012’s count is not given in the sections we annotate, so none is invented here.
Metric & protocolPartial credit, one baseline, fixed decoding
The score is checkpoint-based pass rate. For each trial, the program earns the fraction of the scenario’s checkpoints it passes; a condition’s reported figure is the average over its trials. Complete passes (every checkpoint green) are recorded but treated as a secondary descriptive statistic — the paper’s reasoning is that scenarios differ widely in checkpoint granularity, and binary success would let coarse-grained scenarios dominate the aggregate. Δ values elsewhere on this site are percentage points against the no-guidance baseline within each experiment.
Everything that could confound the comparison is pinned. Two fixed models — Gemini 3.1 Pro Preview and Gemini 3.1 Flash Lite Preview. Low-temperature decoding. A 16,384-token output budget. The task description always enters through the user-facing contents field; the control representation — gene, skill, or nothing — enters separately through systemInstruction. The execution pipeline, timeout policy, and scenario set are shared. The representation is the only moving part.
| Probe | Retained trials | Focus |
|---|---|---|
| Skill Probe | 1,440 | Why documentation packages misfire as control |
| Gene Probe | 1,890 | Construction, robustness, dilution, composition |
| Evolution Probe | 1,260 | Carriers, structure, and failure encoding over time |
| Total | 4,590 | — |
Counts from §4.1–§4.3 of the paper; Appendix B documents the full protocol.
Two honest limits, both stated in the paper itself: the scenarios are scientific code-solving tasks, so transfer to other agent environments is a hypothesis, not a finding; and the CritPt evolution runs are a separate benchmark with its own pairing. Where those caveats matter to a number, the findings page repeats them in place.
9 / 14 checkpoints pass → trial pass rate = 9/14 ≈ 0.643
Why not score pass/fail? Because these scenarios fail partially and informatively. A program can parse the catalog correctly and compute distances correctly, yet still miss the b-value estimate. Checkpoint scoring keeps that signal: each cell above is one sub-task, and a trial’s score is simply the fraction that lit up.
Illustrative trial, not from the paper’s runs. Scenario size matches S026_earthquake_catalog (14 checkpoints).