Arctic benchmark pin

This pins the from-checkout, turnkey Arctic benchmark run by scripts/bench_arctic.py on current main. Its basis is MULTIPRON-ONLY.

The live cells are on/big and on/slt55. The off/big and off/slt55 cells are retained as retired/historical; the off-mode cells are neither trained nor decoded by the current benchmark run.

The machine-readable record, upstream-oracle sidecar, its documented provenance and reconstruction limits, and generated paired analysis preserve the per-utterance rows, live comparisons, and the complete engine, model, corpus, transcript, language-model, dictionary, and decoder identities.

The pin was accepted on a re-derived basis: its measured deltas are the documented baseline. They are not a claim that every cell is at zero delta from the preserved upstream model.

The record’s record_binding is a self-recomputable consistency and tamper-evidence digest over the schema, identities, conditions, models, and measurement rows. It catches an edited field when the digest is left stale, but it does not prove historical provenance: anyone able to relabel the record can recompute it. Actual producing-run provenance rests on the retained resolved configuration, training log, model definition, and artifact hashes, together with the re-derivation gate described below.

Measurement identity

The decode path is a defining condition of this measurement. The live cells are decoded from WAV through pinned PocketSphinx 5.1.1 using Python 3.12.12, native library SHA-256 cee055880a3fde6ae54439b353d0f243246db2a266af5a269cca5890a616ea04, and decode dictionary SHA-256 b9d8271957f978287620d9b20a79e12b0b84470f520942c580145570021d0588. The engine is pstrain 0.3.0 at 80ca415. A result obtained through another decode path is not the same measurement even when the acoustic-model bytes are identical.

The retired off-mode cells were not measured on that path. They were decoded by pstrain 0.1.0 at 740f112 using Python 3.12.12, native library SHA-256 6a5da2377c3b2b033b35d93a12a57bb869413bbc98f045d6c3f3652585792be3, and decode dictionary SHA-256 24ff2852a707b63f499fd968294d5e4c02d44e0eb1ec511e40be1f380d785846. That identity travels with them in the record’s historical_provenance and is never inherited from a later run.

The complete commit and artifact hashes are in the machine-readable record, which is authoritative for both identities.

Pin conditions

Condition

Pinned value

Band

BM1

Language model

SHA-256 2cf11ab0474a0bdd165cbee59db674b05764fdb00bf6f9824c0dccce571637b5

Decode dictionary

SHA-256 b9d8271957f978287620d9b20a79e12b0b84470f520942c580145570021d0588

Filler dictionary

SHA-256 fb50883998c41a5030c2a602965935c647563321e84a86f2adabb377ec24b49c

Shared training

3 states, 200 senones, a_beam=1e-90, b_beam=1e-10, maximum skip fraction 0.05, retry beam factor 1e10, tree state weights [1.0, 0.05, 0.0], ssplitmax=7, ssplitthr=0, csplitmax=2000, csplitthr=0, mwfloor=1e-8, 12 question permutations, 20 questions/state, 1 question iteration

Basis

MULTIPRON-ONLY; off cells retained as retired history

Multipron on training

Frozen settings: multipron_training=true, transcript-reachable untied inventory, optional final silence, one retry at a beam widened by 1e10; CI, untied, and tied 1–10 iterations, convergence 0.001

Acoustic features

16 kHz, 13 cepstra, 25 filters, 512-point FFT, 130–6800 Hz, alpha 0.97, 1s_c_d_dd, lifter 22, DCT, no AGC, batch CMN, no variance normalization

Split

Seed 42, test count 0

Decoder

beam=pbeam=lpbeam=lponlybeam=fwdflatbeam=1e-80; wbeam=fwdflatwbeam=1e-40; pl_window=5, lw=10, wip=0.2

Bootstrap

Matched-pair percentile, 100,000 resamples, seed 7; big cells speaker-stratified

Configuration provenance by result cell

The live cells on/slt55 and on/big come from the named benchmark profile on. There are no run-time overrides. These are the complete differences from the shipped schema defaults; an unlisted setting equals its shipped default. The record’s conditions and each cell’s provenance come from the same resolved build-child snapshot, and validation rejects disagreement between them. Its only semantic difference from shipped product defaults is split.test_count=0, which keeps the established external evaluation cells intact.

The comparison with shipped defaults covers only the configuration blocks the benchmark consumes: those whose every field is read by a training-pipeline stage. The benchmark trains and decodes; it does not run forced alignment. So its alignment settings are frozen in its own configuration, at the values the pinned run resolved (the retry acceptance check, which did not exist then, is frozen off), and they are not compared with shipped defaults. A change to a shipped alignment default therefore does not touch the record.

Cells

Setting

Shipped default

Cell value

Winning source kind

on/slt55, on/big

split.test_count

null

0

project-profile

Provenance correction

The previous record carried its on-mode numbers from a 7f13286 all-triphone run but mislabeled them as a 578f6a9 transcript-reachable run. On 2026-08-17, the pin was remeasured on a Linux x86-64 host at bbb2fef under the declared transcript-reachable profile. The retained run includes its resolved configuration, training log, and an 11,883-row CD-untied mdef. It reproduced the published rows byte-for-byte: on/big SHA-256 11d512dbd1c412bd56a94717b3d91b9fbfdb2ee97c2c169db9c0a9e749e4a977 and on/slt55 SHA-256 f3fa77a1a2cbf51138f0f7c375b8392a9ce707d41a3bba41613cfbb44cdd0d54. The WER results are unchanged; the record now carries the measured engine, configuration, and native-library provenance.

The corpus archive identities are BDL 26b91aaf48b2799b2956792b4632c2f926cd0542f402b5452d5adecb60942904, CLB 3f16dc3f3b97955ea22623efb33b444341013fc660677b2e170efdcc959fa7c6, RMS c6dc11235629c58441c071a7ba8a2d067903dfefbaabc4056d87da35b72ecda4, and SLT 7c173297916acf3cc7fcab2713be4c60b27312316765a90934651d367226b4ea (SHA-256; 1,132 WAVs each). The transcript identities are big 5737c7296d491df39aaa2db24ab03e6acf6e058c10bd7d6b331ffdbf242c5b58, SLT-55 1de4e31c934ea6c5cd414307b8cf4f71c0d846adfd68edf04c4ececaafa4c532, full SLT cce3d341c02445d3aa91453f1d7f0fa3097f6ab6df76264336277ca2c54f3085, and pin training 28788cd1ce2269d344b50420d74007fa8c443778680724f6334e2712ea110959. The 1,043-utterance training fileids hash is 8ce9a55c5929f6f86579ee1b244c38fd4d0a9d41e436e2057337f74c1bb4d631. The JSON record is authoritative for the complete frozen configuration, model-file identities, and resource metadata. Its internal digest does not replace the retained producing-run evidence or establish who produced its measurement rows.

Comparability

The live on/slt55 and on/big arms are comparable as paired decode measurements: both models use the same current engine, dictionary, language model, decoder settings, and scoring path. They are NOT COMPARABLE for implementation attribution because the oracle model’s producing host, architecture, source revision, build, training configuration, inputs, and full lineage are unknown. The retired off/slt55 and off/big arms are comparable as paired historical decode measurements on their own recorded path, NOT COMPARABLE to the live pin because that path and its resources differ, and NOT COMPARABLE for implementation attribution because their model-producing identities are missing or incomplete.

Baseline

Delta is pstrain minus the preserved upstream oracle in WER percentage points. The bands are paired 95% percentile bootstrap summaries. The SLT-55 cells are one speaker and resample its utterances. The big cells resample utterances within each of their three speakers, holding each speaker’s utterance count fixed. That conditions on those three speakers: for a paired delta it gives an interval numerically indistinguishable from an iid utterance bootstrap, and it does not account for speaker-to-speaker variation in the delta.

That variation is present. In the live big cell the per-speaker deltas are +0.67 pp for bdl, -0.54 pp for clb, and +0.52 pp for rms, so they disagree in sign, and their variance is roughly 2.7 times what within-speaker sampling alone would produce. A speaker-level cluster bootstrap over the recorded rows (draw the speakers with replacement, then utterances with replacement within each drawn speaker; 100,000 resamples, seed 7) gives [-0.59, +0.91] for the live big cell, about 65% wider than the tabled interval, and it too straddles zero. The null conclusion therefore holds under either method, but the tabled big-cell bands describe these three voices, which the model was not trained on. With three speakers, no interval supports a claim about unseen voices in general.

Mode

Cell

pstrain WER

Oracle WER

Delta pp

Paired 95% CI

Paired decode

Implementation attribution

Interpretation

off (retired)

SLT-55

28.8499

28.8499

+0.0000

[-4.7059, +4.7619]

COMPARABLE

NOT COMPARABLE

historical only

on

SLT-55

27.0955

28.0702

-0.9747

[-3.6965, +1.6575]

COMPARABLE

NOT COMPARABLE

no statistically significant difference

off (retired)

big

76.6393

74.7915

+1.8478

[+1.3257, +2.3646]

COMPARABLE

NOT COMPARABLE

historical only

on

big

75.7221

75.5053

+0.2168

[-0.2414, +0.6713]

COMPARABLE

NOT COMPARABLE

no statistically significant difference

The live rows come from the record and the resource-matched oracle sidecar through scripts/regenerate_arctic_paired_analysis.py; the retired rows come from the sidecar’s preserved historical comparison. Both are regenerated from the checked-in artifacts, so an amended measurement cannot leave this table stale.

The live oracle rows are re-decoded from the preserved upstream model bytes whenever the pinned decode resources change, so the two arms of the live comparison always consume the same dictionary and language model. The retired rows are the preserved historical comparison and are reported on their own resources; see oracle provenance.

Gap composition

This subsection is a retained earlier experiment. Its figures were computed against the 2026-08-11 oracle decode and the dictionary of that era, so they describe how the gap decomposed then; they are not the live baseline above and do not move with it.

Cross-decoding the preserved earlier-era pstrain models through this pin’s decode path shows that both model instance and decode path contribute. In the big off cell, the preserved model is +1.344 pp behind the oracle (95% CI [+0.832, +1.855]) and is -0.504 pp better than the pin-retrained model (95% CI [-0.958, -0.050]): 72.7% of the pin gap is present in the preserved model on the modern path and 27.3% is the retraining increment. In the big on cell, the corresponding values are +0.787 pp versus oracle (95% CI [+0.253, +1.320]) and -0.427 pp versus the pin model (95% CI [-0.881, +0.027]), or 64.8% and 35.2% of the pin gap respectively.

Across eras, moving the preserved pstrain model to the pin path changes WER by only +0.070 pp off and +0.157 pp on, while the byte-identical upstream oracle improves by 0.597 pp off and 0.504 pp on. The decode path is therefore model-sensitive; the differential decode-path response supplies most of the era-to-era delta movement, and pin retraining supplies a further roughly 0.43–0.50 pp.

Forward gate

Future runs compare matched pairs against the pinned per-utterance rows. For each live cell the gate computes the paired 95% bootstrap interval of the current run minus the pinned rows, in WER percentage points, and passes only when the upper end of that interval is at or below zero. That is a non-inferiority test with a zero margin: a run passes when it reproduces the pinned rows exactly, or when its interval lies entirely at or below zero so the data rule out any worsening. An interval that straddles zero fails, even when the point estimate favors the new run, because that change cannot be told apart from a regression. The gate never absorbs a change to the recorded results: moving them is always a deliberate re-pin. The live cells’ standing against the preserved upstream models is part of this documented baseline and is not itself a regression; the retired cells’ larger gap is history and is not a target.

Both arms of the live comparison must stay on the same decode resources. When the pinned dictionary or language model changes, re-decode the preserved oracle models with scripts/regenerate_arctic_oracle.py before regenerating the paired analysis; make pin-check refuses arms that name different resource bytes.

Decode-path transport

Earlier program measurements used precomputed-feature batch decoding rather than this WAV decode path. Those measurements remain bound to that path: a byte-identical acoustic model shifts by 0.50–0.60 pp WER between the two paths. This pin consequently records its full decode identity. Future comparisons must either hold that identity fixed or cross-decode the compared models.

Training skips and decode coverage

The live pin has no accepted training skips. In particular, arctic_a0587 trains under product defaults and neither former exception fires. The training.accept_arctic_a0587_known_skip knob is retained solely to describe the retired off profile’s provenance; that profile carries the true value, while the live on profile disables it (and the a0302 exception band). Thus all live cells run exception-free.

A live cell must decode every utterance its denominator names. A cell’s absolute WER is summed from its per-utterance rows, and those rows are the decoded subset, so a live cell whose decoded fell below its decode_denominator would publish a headline rate over a denominator its own coverage fields contradict, and a comparison against a record carrying the same shortfall would report no difference at all. validate_record therefore refuses to pin such a cell, which places the refusal ahead of record emission, fresh-record adoption, and make config-check alike. The retired off cells are not re-gated: they were measured on a path that no longer exists, and their coverage is a fact about the past.

The paired comparison stays ungated, because it is sound under a shortfall: paired_delta_ci requires identical utterance sets, so lost coverage cannot bias the delta between the arms. A run that transiently decodes fewer utterances than the record still compares, and drift from the pinned decoded/denominator is surfaced in field_differences beside the row for deliberate record adoption rather than raising or failing the WER gate. Such a run simply cannot become the pin.

Condition contract maintenance

The record owns the condition fields it contains. A changed or removed pinned value is fatal; a live field added after the record was written is reported as uncovered but does not masquerade as benchmark drift. make config-check runs this authentication in the ordinary suite. Adopt newly added fields deliberately, without changing existing pins or any measured result, with python scripts/check_arctic_pin.py --adopt-uncovered.

Fresh-record adoption also requires exact full-cell equality for both retired off cells. Every serialized field in each retired cell—including bootstrap summaries and any future field—is stable; any addition, removal, or value change is refused.

Replicability

The benchmark uses canonical ordering and is deterministic. Native Mach-O fingerprints omit randomized UUID and derived code-signature bytes, and model identity covers every acoustic-model artifact while excluding timing, provenance, and completion metadata. The same checkout and resource band therefore produce the same pinned content identities; environment identity is recorded separately.