# Arctic benchmark pin This pins the from-checkout, turnkey Arctic benchmark run by `scripts/bench_arctic.py` on current main. Its basis is `MULTIPRON-ONLY`. The live cells are `on/big` and `on/slt55`. The `off/big` and `off/slt55` cells are retained as `retired/historical`; the off-mode cells are neither trained nor decoded by the current benchmark run. The machine-readable [record](../../evidence/arctic-pin/record.json), upstream-oracle [sidecar](../../evidence/arctic-pin/oracle-sidecar.json), its documented [provenance and reconstruction limits](oracle-provenance.md), and generated [paired analysis](../../evidence/arctic-pin/paired-analysis.json) preserve the per-utterance rows, live comparisons, and the complete engine, model, corpus, transcript, language-model, dictionary, and decoder identities. The pin was accepted on a re-derived basis: its measured deltas are the documented baseline. They are not a claim that every cell is at zero delta from the preserved upstream model. The record's `record_binding` is a self-recomputable consistency and tamper-evidence digest over the schema, identities, conditions, models, and measurement rows. It catches an edited field when the digest is left stale, but it does not prove historical provenance: anyone able to relabel the record can recompute it. Actual producing-run provenance rests on the retained resolved configuration, training log, model definition, and artifact hashes, together with the re-derivation gate described below. ## Measurement identity The decode path is a defining condition of this measurement. The live cells are decoded from WAV through pinned PocketSphinx 5.1.1 using Python 3.12.12, native library SHA-256 `cee055880a3fde6ae54439b353d0f243246db2a266af5a269cca5890a616ea04`, and decode dictionary SHA-256 `b9d8271957f978287620d9b20a79e12b0b84470f520942c580145570021d0588`. The engine is pstrain 0.3.0 at `80ca415`. A result obtained through another decode path is not the same measurement even when the acoustic-model bytes are identical. The retired off-mode cells were not measured on that path. They were decoded by pstrain 0.1.0 at `740f112` using Python 3.12.12, native library SHA-256 `6a5da2377c3b2b033b35d93a12a57bb869413bbc98f045d6c3f3652585792be3`, and decode dictionary SHA-256 `24ff2852a707b63f499fd968294d5e4c02d44e0eb1ec511e40be1f380d785846`. That identity travels with them in the record's `historical_provenance` and is never inherited from a later run. The complete commit and artifact hashes are in the machine-readable record, which is authoritative for both identities. ## Pin conditions | Condition | Pinned value | |---|---| | Band | BM1 | | Language model | SHA-256 `2cf11ab0474a0bdd165cbee59db674b05764fdb00bf6f9824c0dccce571637b5` | | Decode dictionary | SHA-256 `b9d8271957f978287620d9b20a79e12b0b84470f520942c580145570021d0588` | | Filler dictionary | SHA-256 `fb50883998c41a5030c2a602965935c647563321e84a86f2adabb377ec24b49c` | | Shared training | 3 states, 200 senones, `a_beam=1e-90`, `b_beam=1e-10`, maximum skip fraction 0.05, retry beam factor `1e10`, tree state weights `[1.0, 0.05, 0.0]`, `ssplitmax=7`, `ssplitthr=0`, `csplitmax=2000`, `csplitthr=0`, `mwfloor=1e-8`, 12 question permutations, 20 questions/state, 1 question iteration | | Basis | `MULTIPRON-ONLY`; off cells retained as retired history | | Multipron on training | Frozen settings: `multipron_training=true`, transcript-reachable untied inventory, optional final silence, one retry at a beam widened by `1e10`; CI, untied, and tied 1–10 iterations, convergence 0.001 | | Acoustic features | 16 kHz, 13 cepstra, 25 filters, 512-point FFT, 130–6800 Hz, alpha 0.97, `1s_c_d_dd`, lifter 22, DCT, no AGC, batch CMN, no variance normalization | | Split | Seed 42, test count 0 | | Decoder | `beam=pbeam=lpbeam=lponlybeam=fwdflatbeam=1e-80`; `wbeam=fwdflatwbeam=1e-40`; `pl_window=5`, `lw=10`, `wip=0.2` | | Bootstrap | Matched-pair percentile, 100,000 resamples, seed 7; big cells speaker-stratified | ### Configuration provenance by result cell The live cells `on/slt55` and `on/big` come from the named benchmark profile `on`. There are no run-time overrides. These are the complete differences from the shipped schema defaults; an unlisted setting equals its shipped default. The record's conditions and each cell's provenance come from the same resolved build-child snapshot, and validation rejects disagreement between them. Its only semantic difference from shipped product defaults is `split.test_count=0`, which keeps the established external evaluation cells intact. The comparison with shipped defaults covers only the configuration blocks the benchmark consumes: those whose every field is read by a training-pipeline stage. The benchmark trains and decodes; it does not run forced alignment. So its alignment settings are frozen in its own configuration, at the values the pinned run resolved (the retry acceptance check, which did not exist then, is frozen off), and they are not compared with shipped defaults. A change to a shipped alignment default therefore does not touch the record. | Cells | Setting | Shipped default | Cell value | Winning source kind | |---|---|---:|---:|---| | on/slt55, on/big | `split.test_count` | `null` | `0` | `project-profile` | ## Provenance correction The previous record carried its on-mode numbers from a `7f13286` all-triphone run but mislabeled them as a `578f6a9` transcript-reachable run. On 2026-08-17, the pin was remeasured on a Linux x86-64 host at `bbb2fef` under the declared transcript-reachable profile. The retained run includes its resolved configuration, training log, and an 11,883-row CD-untied mdef. It reproduced the published rows byte-for-byte: `on/big` SHA-256 `11d512dbd1c412bd56a94717b3d91b9fbfdb2ee97c2c169db9c0a9e749e4a977` and `on/slt55` SHA-256 `f3fa77a1a2cbf51138f0f7c375b8392a9ce707d41a3bba41613cfbb44cdd0d54`. The WER results are unchanged; the record now carries the measured engine, configuration, and native-library provenance. The corpus archive identities are BDL `26b91aaf48b2799b2956792b4632c2f926cd0542f402b5452d5adecb60942904`, CLB `3f16dc3f3b97955ea22623efb33b444341013fc660677b2e170efdcc959fa7c6`, RMS `c6dc11235629c58441c071a7ba8a2d067903dfefbaabc4056d87da35b72ecda4`, and SLT `7c173297916acf3cc7fcab2713be4c60b27312316765a90934651d367226b4ea` (SHA-256; 1,132 WAVs each). The transcript identities are big `5737c7296d491df39aaa2db24ab03e6acf6e058c10bd7d6b331ffdbf242c5b58`, SLT-55 `1de4e31c934ea6c5cd414307b8cf4f71c0d846adfd68edf04c4ececaafa4c532`, full SLT `cce3d341c02445d3aa91453f1d7f0fa3097f6ab6df76264336277ca2c54f3085`, and pin training `28788cd1ce2269d344b50420d74007fa8c443778680724f6334e2712ea110959`. The 1,043-utterance training fileids hash is `8ce9a55c5929f6f86579ee1b244c38fd4d0a9d41e436e2057337f74c1bb4d631`. The JSON record is authoritative for the complete frozen configuration, model-file identities, and resource metadata. Its internal digest does not replace the retained producing-run evidence or establish who produced its measurement rows. ## Comparability The live `on/slt55` and `on/big` arms are comparable as paired decode measurements: both models use the same current engine, dictionary, language model, decoder settings, and scoring path. They are **NOT COMPARABLE** for implementation attribution because the oracle model's producing host, architecture, source revision, build, training configuration, inputs, and full lineage are unknown. The retired `off/slt55` and `off/big` arms are comparable as paired historical decode measurements on their own recorded path, **NOT COMPARABLE** to the live pin because that path and its resources differ, and **NOT COMPARABLE** for implementation attribution because their model-producing identities are missing or incomplete. ## Baseline Delta is pstrain minus the preserved upstream oracle in WER percentage points. The bands are paired 95% percentile bootstrap summaries. The SLT-55 cells are one speaker and resample its utterances. The big cells resample utterances within each of their three speakers, holding each speaker's utterance count fixed. That conditions on those three speakers: for a paired delta it gives an interval numerically indistinguishable from an iid utterance bootstrap, and it does not account for speaker-to-speaker variation in the delta. That variation is present. In the live big cell the per-speaker deltas are +0.67 pp for bdl, -0.54 pp for clb, and +0.52 pp for rms, so they disagree in sign, and their variance is roughly 2.7 times what within-speaker sampling alone would produce. A speaker-level cluster bootstrap over the recorded rows (draw the speakers with replacement, then utterances with replacement within each drawn speaker; 100,000 resamples, seed 7) gives [-0.59, +0.91] for the live big cell, about 65% wider than the tabled interval, and it too straddles zero. The null conclusion therefore holds under either method, but the tabled big-cell bands describe these three voices, which the model was not trained on. With three speakers, no interval supports a claim about unseen voices in general. | Mode | Cell | pstrain WER | Oracle WER | Delta pp | Paired 95% CI | Paired decode | Implementation attribution | Interpretation | |---|---|---:|---:|---:|---:|---|---|---| | off (retired) | SLT-55 | 28.8499 | 28.8499 | +0.0000 | [-4.7059, +4.7619] | COMPARABLE | NOT COMPARABLE | historical only | | on | SLT-55 | 27.0955 | 28.0702 | -0.9747 | [-3.6965, +1.6575] | COMPARABLE | NOT COMPARABLE | no statistically significant difference | | off (retired) | big | 76.6393 | 74.7915 | +1.8478 | [+1.3257, +2.3646] | COMPARABLE | NOT COMPARABLE | historical only | | on | big | 75.7221 | 75.5053 | +0.2168 | [-0.2414, +0.6713] | COMPARABLE | NOT COMPARABLE | no statistically significant difference | The live rows come from the record and the resource-matched oracle sidecar through `scripts/regenerate_arctic_paired_analysis.py`; the retired rows come from the sidecar's preserved historical comparison. Both are regenerated from the checked-in artifacts, so an amended measurement cannot leave this table stale. The live oracle rows are re-decoded from the preserved upstream model bytes whenever the pinned decode resources change, so the two arms of the live comparison always consume the same dictionary and language model. The retired rows are the preserved historical comparison and are reported on their own resources; see [oracle provenance](oracle-provenance.md). ### Gap composition This subsection is a retained earlier experiment. Its figures were computed against the 2026-08-11 oracle decode and the dictionary of that era, so they describe how the gap decomposed then; they are not the live baseline above and do not move with it. Cross-decoding the preserved earlier-era pstrain models through this pin's decode path shows that both model instance and decode path contribute. In the big off cell, the preserved model is +1.344 pp behind the oracle (95% CI [+0.832, +1.855]) and is -0.504 pp better than the pin-retrained model (95% CI [-0.958, -0.050]): 72.7% of the pin gap is present in the preserved model on the modern path and 27.3% is the retraining increment. In the big on cell, the corresponding values are +0.787 pp versus oracle (95% CI [+0.253, +1.320]) and -0.427 pp versus the pin model (95% CI [-0.881, +0.027]), or 64.8% and 35.2% of the pin gap respectively. Across eras, moving the preserved pstrain model to the pin path changes WER by only +0.070 pp off and +0.157 pp on, while the byte-identical upstream oracle improves by 0.597 pp off and 0.504 pp on. The decode path is therefore model-sensitive; the differential decode-path response supplies most of the era-to-era delta movement, and pin retraining supplies a further roughly 0.43–0.50 pp. ## Forward gate Future runs compare matched pairs against the pinned per-utterance rows. For each live cell the gate computes the paired 95% bootstrap interval of the current run minus the pinned rows, in WER percentage points, and passes only when the upper end of that interval is at or below zero. That is a non-inferiority test with a zero margin: a run passes when it reproduces the pinned rows exactly, or when its interval lies entirely at or below zero so the data rule out any worsening. An interval that straddles zero fails, even when the point estimate favors the new run, because that change cannot be told apart from a regression. The gate never absorbs a change to the recorded results: moving them is always a deliberate re-pin. The live cells' standing against the preserved upstream models is part of this documented baseline and is not itself a regression; the retired cells' larger gap is history and is not a target. Both arms of the live comparison must stay on the same decode resources. When the pinned dictionary or language model changes, re-decode the preserved oracle models with `scripts/regenerate_arctic_oracle.py` before regenerating the paired analysis; `make pin-check` refuses arms that name different resource bytes. ## Decode-path transport Earlier program measurements used precomputed-feature batch decoding rather than this WAV decode path. Those measurements remain bound to that path: a byte-identical acoustic model shifts by 0.50–0.60 pp WER between the two paths. This pin consequently records its full decode identity. Future comparisons must either hold that identity fixed or cross-decode the compared models. ## Training skips and decode coverage The live pin has no accepted training skips. In particular, `arctic_a0587` trains under product defaults and neither former exception fires. The `training.accept_arctic_a0587_known_skip` knob is retained solely to describe the retired `off` profile's provenance; that profile carries the true value, while the live `on` profile disables it (and the a0302 exception band). Thus all live cells run exception-free. A live cell must decode every utterance its denominator names. A cell's absolute WER is summed from its per-utterance rows, and those rows are the decoded subset, so a live cell whose `decoded` fell below its `decode_denominator` would publish a headline rate over a denominator its own coverage fields contradict, and a comparison against a record carrying the same shortfall would report no difference at all. `validate_record` therefore refuses to pin such a cell, which places the refusal ahead of record emission, fresh-record adoption, and `make config-check` alike. The retired `off` cells are not re-gated: they were measured on a path that no longer exists, and their coverage is a fact about the past. The paired comparison stays ungated, because it is sound under a shortfall: `paired_delta_ci` requires identical utterance sets, so lost coverage cannot bias the delta between the arms. A run that transiently decodes fewer utterances than the record still compares, and drift from the pinned `decoded/denominator` is surfaced in `field_differences` beside the row for deliberate record adoption rather than raising or failing the WER gate. Such a run simply cannot become the pin. ## Condition contract maintenance The record owns the condition fields it contains. A changed or removed pinned value is fatal; a live field added after the record was written is reported as uncovered but does not masquerade as benchmark drift. `make config-check` runs this authentication in the ordinary suite. Adopt newly added fields deliberately, without changing existing pins or any measured result, with `python scripts/check_arctic_pin.py --adopt-uncovered`. Fresh-record adoption also requires exact full-cell equality for both retired `off` cells. Every serialized field in each retired cell—including bootstrap summaries and any future field—is stable; any addition, removal, or value change is refused. ## Replicability The benchmark uses canonical ordering and is deterministic. Native Mach-O fingerprints omit randomized UUID and derived code-signature bytes, and model identity covers every acoustic-model artifact while excluding timing, provenance, and completion metadata. The same checkout and resource band therefore produce the same pinned content identities; environment identity is recorded separately.