RESEARCH RELEASE · ONE-SHOT, RESTRICTED TASK · READ THE FAILURE BOUNDARIES
A build story · nano-agent · code-speed
1.5B base + LoRAone proposal per case24 development + two 30-case frozen tracks

The optimization that forgot the question

A small model passed 23 of 24 code-speed development tasks. Then an apparently sensible data update cut the result almost in half. An independent verifier showed exactly what went wrong—and the frozen tests showed what the repair still could not do.

01 / The hookThe patch that changed the answer

Suppose a Python function counts distinct values after normalization. A fluent rewrite builds a frequency table and counts values seen exactly once. It looks efficient. On [4, 4, 9], though, the original question has answer 2; the rewrite returns 1.

# Conceptual illustration of the observed failure class
distinct = len(set(xs))                 # [4, 4, 9] → 2
singletons = sum(n == 1 for n in counts.values())  # → 1

The actual development case also performed modulo normalization, making the invariant easier to miss. This is why a score for “a patch was emitted” would be misleading. A performance improvement is useful only after behavior is preserved.

02 / The systemThe verification boundary

We trained a PEFT LoRA adapter on Qwen/Qwen2.5-Coder-1.5B-Instruct to emit SEARCH/REPLACE patches for a restricted pure-Python function. The production Rust controller applies one greedy proposal. A network-isolated verifier compares the result with the original on hidden edge, random, and benchmark inputs. Only a correct candidate with lower Cachegrind instruction references counts as success.

Figure 1 · one-shot patch-and-measure path
Slow function
model-visible
→
1.5B adapter
one patch
→
Rust apply
step
→
Hidden-input
correctness
→
Cachegrind
Ir reward
→
Accept only
if both pass
Model-facing context excludes hidden inputs, private seeds, generator metadata, and known-fast source. The diagram is a protocol, not a claim of formal equivalence.

The verifier runs a deliberately small language subset with Bubblewrap isolation and resource limits. Imports, network and filesystem APIs, introspection, decorators, and nested functions are outside scope. Its reward, for a correct repeatable candidate, is log(reference_Ir / candidate_Ir); wrong outputs receive zero. Ir is the simulated instruction-reference count of the whole Python process, not wall-clock latency. Fixed interpreter overhead can dilute apparent speedups.

Important boundary

Finite hidden tests can reject observed mistakes. They cannot prove that two programs are equivalent on every possible input. This is a specialist proposer inside a verifier-gated loop, not a general optimizer or a public code-execution service.

03 / The buildHow the training data was built

The recipe begins with deterministic single-function optimization problems. Each row identifies its transformation family and template; the trusted parent verifies its known-fast pair but never writes that known-fast source into the model-facing JSONL. The v1 split contains 152 train, 38 development, 30 frozen seen-family, and 30 frozen family-heldout cases. A separate 24-case behavioral development set drives the later 1.5B checkpoint decisions.

  1. Verify candidate problems. The same correctness and Cachegrind gate checks generated slow/fast pairs. An unverified staging corpus is not training data.
  2. Collect teacher trajectories. A frontier teacher generated candidate trajectories. Only teacher turns that passed the real controller/verifier were retained. v6.1 reused these traces without new teacher calls.
  3. Build the SFT mixture. v5 combined 530 verified teacher turns, 502 focused extra copies across six training-only invariant families, and 115 v1 carryover rows. v6 used 550 teacher turns and one repair trace but no focused copies. v6.1 retained the new turns and restored 504 focused copies, for 1,170 rows.
  4. Protect the split boundary. A leakage audit checked case IDs, templates, and normalized source hashes against development and frozen cases. The training job split by case ID before replay weighting, so duplicate examples could not cross into its internal validation partition.
  5. Train and measure behavior. Both release variants used LoRA rank 16, alpha 32, dropout 0.05, one epoch, learning rate 5e-5, batch 4, gradient accumulation 4, seed 7, and 40% effective v1 replay. The development test used the unchanged controller and verifier.
Integrity gates before v6.1 GPU training

The 1,170-row contract probe checked assistant supervision (116–394 tokens), the leakage audit was clean, and a one-step CPU smoke completed. These gates verified the input and training contract; none was counted as a behavioral win.

04 / The regressionMore verified data, fewer correct patches

v5 passed 23/24 behavioral development tasks. v6 added verified teacher coverage but omitted the 502 concentrated copies that had reinforced semantic invariants. It passed only 12/24. Nine misses were output mismatches and three were candidate errors; all twelve were classified as semantic, with no formatting, patch-application, or verifier-infrastructure failures. v6.1 kept the new coverage, restored focus, and reached 21/24.

Figure 2 · development successes / 24Development success across v5, v6, and v6.1v5 passed 23 of 24, v6 passed 12 of 24, v6.1 passed 21 of 24.24120231221v5v6v6.1
Same 24-case behavioral development set, one greedy proposal per case. Counts are point estimates from one training run per configuration; this is not a controlled causal ablation.

The development failures were not evenly spread. v6 dropped from 3/3 to 0/3 on both complement lookup and running aggregate, and from 2/3 to 0/3 on fixed sliding windows. v6.1 restored 3/3 on complement lookup and running aggregate, but still missed one distinct-cardinality, one stable-deduplication, and one sliding-window case. The data-mix change is a plausible explanation, not proof that duplication alone caused the fall and recovery.

VariantCorpus rowsInternal eval lossBehavioral dev
v51,1470.2252523/24
v66660.2585812/24
v6.11,1700.1823721/24

Internal loss is not an apples-to-apples behavioral score: the corpora and validation partitions differ. Its role here is to show why training fit alone was not the promotion gate.

05 / The evidenceAnatomy of a semantic miss

Three v6.1 development misses demonstrate distinct invariant failures. One counted singletons instead of distinct normalized values. One filtered all repeated values rather than keeping each first occurrence—and mixed unnormalized counting with normalized output. One correctly formed a running window but chose the minimum where the reference selected the maximum. The verifier rejected each after applying the patch.

Fresh three-case diagnostic: 0/3

In every v6.1 residue-count-index case, the candidate omitted the reference function’s local bucket_size assignment. The remaining code referenced that now-undefined name; one candidate also reversed query order. These were candidate_error statuses grouped by the evaluator as semantic failures. v5 was not measured on this split.

# Schematic of the omitted dependency; not a verbatim hidden test
def answer(xs, queries):
    bucket_size = 16             # required local fact
    counts = build_counts(xs, bucket_size)
    return [counts.get(q, 0) for q in queries]

# If a rewrite removes the assignment but keeps `bucket_size`,
# it may look structurally faster and still fail at runtime.

Three cases are far too few to estimate general transfer. They are enough to locate a concrete failure mechanism: the model was willing to perform a structural rewrite without reliably carrying local dependencies through it.

06 / The testThe frozen result changes the ending

v5 remained the development-selected checkpoint. We measured v6.1 on frozen tracks afterward for a publication comparison; those scores were not used to select another checkpoint or train to the exact cases. The seen-family track asks whether learned transformations survive new templates. The family-heldout track asks whether the model can handle whole transformation families excluded from training.

Frozen trackv5v6.1What changed?
Seen-family24/3024/30Equal totals, different case-level errors
Family-heldout14/3016/30+1 string concatenation, +1 sort selection

The equal seen-family total hides a trade: v6.1 gained one complement-lookup and one fixed-window case, but lost one distinct-cardinality and one stable-deduplication case. Of its six seen-family misses, five changed semantics and one was correct but not faster.

Figure 3 · held-out transformation families / 10 eachFrozen family-heldout resultsString concatenation: v5 7 of 10, v6.1 8 of 10. Sort selection: 7 and 8. Indexed lookup: both zero.1050787800String concat.Sort selectionIndexed lookupv5v6.1
Two 30-case frozen tracks must not be pooled. This figure shows only the family-heldout track, where each family has ten cases. Both adapters failed every indexed-lookup case.

Thirteen of v6.1’s fourteen family-heldout misses were semantic failures; one was correct but not faster. The two-case aggregate gain is real for these cases, but it is not evidence of broad optimization ability. Indexed lookup remained a complete failure for both variants.

07 / The audit trailHow to check what we actually measured

Each result report names its split, one-shot protocol, failure categories, and checksums for retained raw JSON. The claim–evidence matrix ties public claims to those reports and lists claims we deliberately exclude. Review the v5 result, v6 regression, v6.1 development result, and v6.1 frozen result alongside the charts.

v6.1 reportFull JSON SHA-256
Frozen seen-family00a41526db5fac9e963e6c41a4bd59a1888ba2fcd3ce2e2266a7d384f65b7dd2
Frozen family-heldoutadc0066bde045e9622654de6dca8bc16d2c7c42ea30f9d2257eb64e5a09ec383
Fresh three-case diagnostic05a66dab3d26b7fbe2266e41eb1913b5faac74d59df5fb68fe4e08874d68b1c1

The GCP training archives were copied and SHA-256 verified before VM and boot-disk deletion. The recorded on-demand L4 quote was $0.856856553/hour under a $1.75 two-hour authorization ceiling; actual billing remains pending. The temporary OCI evaluator and model server were stopped after the reports were extracted.

What is intentionally private

Hidden test inputs, private seeds, teacher prompt traces, and training JSONL are not part of the public blog. The source pipeline, split policy, model config, report hashes, and aggregate/case-level safe summaries are the reproducible public boundary.

08 / The learningWhat this taught us

First: the verifier defines the result. Without it, a fluent patch that counts singletons or drops a local constant could be called a speedup. With it, those patches become evidence about where the model is brittle.

Second: more verified data is not enough by itself. v6 added coverage and lost behavioral success. The mix of invariants being reinforced mattered in this sequence, even though one run per configuration cannot identify a sole cause.

Third: token fit is not program behavior. v6.1’s internal evaluation loss was lower than v5’s, but v5 kept the better development score. The thing we want is not simply a more likely patch; it is a patch that preserves the program and earns a measured reward.

Fourth: transfer needs a new boundary. A next experiment should explicitly teach preservation of local dependencies and study indexed lookup using new, truly disjoint evaluation cases. Tuning to the current frozen problems would spend their evidential value. Multiple seeds and paired case-level comparisons would strengthen any future data-mixture conclusion.

Equal release visibility is not equal performance. v5 leads on development; v6.1 ties seen-family, gains two heldout cases, and fails the fresh diagnostic.

09 / The handoffExplore the code and the limits

The repository’s recipes/code-speedup/ directory contains the Rust controller, Python bridge, deterministic corpus generator, leakage checker, SFT code, and evaluation runner. On a compatible Linux host with Bubblewrap, Valgrind, and the required namespaces, the local verifier and data-contract checks start here:

# From the nano-agent repository root
python3 -m pytest recipes/code-speedup/harness -v
python3 -m ruff check recipes/code-speedup
cargo test --workspace

Reproducing the published model scores also requires the exact adapter, pinned verifier image, frozen split artifacts, and model-serving setup described in the release manifest. Do not substitute the command above for a score reproduction: it checks implementation behavior, not the frozen one-shot model run. Private verifier inputs are not distributed.

The release, honestly framed

Two variants. One boundary.

v5 and v6.1 are equal experiment variants in a separate code-speed Hugging Face repository, without a default adapter. Read the paper, its Zenodo record (DOI 10.5281/zenodo.22922767), and the source repository alongside the model card and checksummed evidence: each has different strengths and failures. This is a specialist proposer in a verified loop—not a general code optimizer. The open research question is how to teach it to carry quiet local facts through a structural rewrite, so that the faster program is still the same program.