Arctic benchmark pin
This pins the from-checkout, turnkey Arctic benchmark run by
scripts/bench_arctic.py on current main. Its basis is MULTIPRON-ONLY.
The live cells are on/big and on/slt55. The off/big and off/slt55 cells are retained as
retired/historical; the off-mode cells are neither trained nor decoded by the
current benchmark run.
The machine-readable record, upstream-oracle sidecar, its documented provenance and reconstruction limits, and generated paired analysis preserve the per-utterance rows, live comparisons, and the complete engine, model, corpus, transcript, language-model, dictionary, and decoder identities.
The pin was accepted on a re-derived basis: its measured deltas are the documented baseline. They are not a claim that every cell is at zero delta from the preserved upstream model.
The record’s record_binding is a self-recomputable consistency and
tamper-evidence digest over the schema, identities, conditions, models, and
measurement rows. It catches an edited field when the digest is left stale,
but it does not prove historical provenance: anyone able to relabel the record
can recompute it. Actual producing-run provenance rests on the retained
resolved configuration, training log, model definition, and artifact hashes,
together with the re-derivation gate described below.
Measurement identity
The decode path is a defining condition of this measurement. The live cells are decoded
from WAV through pinned PocketSphinx 5.1.1 using Python 3.12.12, native library SHA-256
cee055880a3fde6ae54439b353d0f243246db2a266af5a269cca5890a616ea04, and decode
dictionary SHA-256 b9d8271957f978287620d9b20a79e12b0b84470f520942c580145570021d0588.
The engine is pstrain 0.3.0 at 80ca415. A result obtained through another decode path
is not the same measurement even when the acoustic-model bytes are identical.
The retired off-mode cells were not measured on that path. They were decoded by pstrain
0.1.0 at 740f112 using Python 3.12.12, native library SHA-256
6a5da2377c3b2b033b35d93a12a57bb869413bbc98f045d6c3f3652585792be3, and decode
dictionary SHA-256 24ff2852a707b63f499fd968294d5e4c02d44e0eb1ec511e40be1f380d785846.
That identity travels with them in the record’s historical_provenance and is never
inherited from a later run.
The complete commit and artifact hashes are in the machine-readable record, which is authoritative for both identities.
Pin conditions
Condition |
Pinned value |
|---|---|
Band |
BM1 |
Language model |
SHA-256 |
Decode dictionary |
SHA-256 |
Filler dictionary |
SHA-256 |
Shared training |
3 states, 200 senones, |
Basis |
|
Multipron on training |
Frozen settings: |
Acoustic features |
16 kHz, 13 cepstra, 25 filters, 512-point FFT, 130–6800 Hz, alpha 0.97, |
Split |
Seed 42, test count 0 |
Decoder |
|
Bootstrap |
Matched-pair percentile, 100,000 resamples, seed 7; big cells speaker-stratified |
Configuration provenance by result cell
The live cells on/slt55 and on/big come from the named benchmark profile
on. There are no run-time overrides. These are the complete differences from
the shipped schema defaults; an unlisted setting equals its shipped default.
The record’s
conditions and each cell’s provenance come from the same resolved build-child
snapshot, and validation rejects disagreement between them. Its only semantic
difference from shipped product defaults is split.test_count=0, which keeps
the established external evaluation cells intact.
The comparison with shipped defaults covers only the configuration blocks the benchmark consumes: those whose every field is read by a training-pipeline stage. The benchmark trains and decodes; it does not run forced alignment. So its alignment settings are frozen in its own configuration, at the values the pinned run resolved (the retry acceptance check, which did not exist then, is frozen off), and they are not compared with shipped defaults. A change to a shipped alignment default therefore does not touch the record.
Cells |
Setting |
Shipped default |
Cell value |
Winning source kind |
|---|---|---|---|---|
on/slt55, on/big |
|
|
|
|
Provenance correction
The previous record carried its on-mode numbers from a 7f13286
all-triphone run but mislabeled them as a 578f6a9 transcript-reachable run.
On 2026-08-17, the pin was remeasured on a Linux x86-64 host at bbb2fef under the declared
transcript-reachable profile. The retained run includes its resolved
configuration, training log, and an 11,883-row CD-untied mdef. It reproduced
the published rows byte-for-byte: on/big SHA-256
11d512dbd1c412bd56a94717b3d91b9fbfdb2ee97c2c169db9c0a9e749e4a977
and on/slt55 SHA-256
f3fa77a1a2cbf51138f0f7c375b8392a9ce707d41a3bba41613cfbb44cdd0d54.
The WER results are unchanged; the record now carries the measured engine,
configuration, and native-library provenance.
The corpus archive identities are BDL
26b91aaf48b2799b2956792b4632c2f926cd0542f402b5452d5adecb60942904,
CLB 3f16dc3f3b97955ea22623efb33b444341013fc660677b2e170efdcc959fa7c6,
RMS c6dc11235629c58441c071a7ba8a2d067903dfefbaabc4056d87da35b72ecda4,
and SLT 7c173297916acf3cc7fcab2713be4c60b27312316765a90934651d367226b4ea
(SHA-256; 1,132 WAVs each). The transcript identities are big
5737c7296d491df39aaa2db24ab03e6acf6e058c10bd7d6b331ffdbf242c5b58,
SLT-55 1de4e31c934ea6c5cd414307b8cf4f71c0d846adfd68edf04c4ececaafa4c532,
full SLT cce3d341c02445d3aa91453f1d7f0fa3097f6ab6df76264336277ca2c54f3085,
and pin training
28788cd1ce2269d344b50420d74007fa8c443778680724f6334e2712ea110959.
The 1,043-utterance training fileids hash is
8ce9a55c5929f6f86579ee1b244c38fd4d0a9d41e436e2057337f74c1bb4d631.
The JSON record is authoritative for the complete frozen configuration,
model-file identities, and resource metadata. Its internal digest does not
replace the retained producing-run evidence or establish who produced its
measurement rows.
Comparability
The live on/slt55 and on/big arms are comparable as paired decode
measurements: both models use the same current engine, dictionary, language
model, decoder settings, and scoring path. They are NOT COMPARABLE for
implementation attribution because the oracle model’s producing host,
architecture, source revision, build, training configuration, inputs, and full
lineage are unknown. The retired off/slt55 and off/big arms are comparable
as paired historical decode measurements on their own recorded path, NOT
COMPARABLE to the live pin because that path and its resources differ, and
NOT COMPARABLE for implementation attribution because their model-producing
identities are missing or incomplete.
Baseline
Delta is pstrain minus the preserved upstream oracle in WER percentage points. The bands are paired 95% percentile bootstrap summaries. The SLT-55 cells are one speaker and resample its utterances. The big cells resample utterances within each of their three speakers, holding each speaker’s utterance count fixed. That conditions on those three speakers: for a paired delta it gives an interval numerically indistinguishable from an iid utterance bootstrap, and it does not account for speaker-to-speaker variation in the delta.
That variation is present. In the live big cell the per-speaker deltas are +0.67 pp for bdl, -0.54 pp for clb, and +0.52 pp for rms, so they disagree in sign, and their variance is roughly 2.7 times what within-speaker sampling alone would produce. A speaker-level cluster bootstrap over the recorded rows (draw the speakers with replacement, then utterances with replacement within each drawn speaker; 100,000 resamples, seed 7) gives [-0.59, +0.91] for the live big cell, about 65% wider than the tabled interval, and it too straddles zero. The null conclusion therefore holds under either method, but the tabled big-cell bands describe these three voices, which the model was not trained on. With three speakers, no interval supports a claim about unseen voices in general.
Mode |
Cell |
pstrain WER |
Oracle WER |
Delta pp |
Paired 95% CI |
Paired decode |
Implementation attribution |
Interpretation |
|---|---|---|---|---|---|---|---|---|
off (retired) |
SLT-55 |
28.8499 |
28.8499 |
+0.0000 |
[-4.7059, +4.7619] |
COMPARABLE |
NOT COMPARABLE |
historical only |
on |
SLT-55 |
27.0955 |
28.0702 |
-0.9747 |
[-3.6965, +1.6575] |
COMPARABLE |
NOT COMPARABLE |
no statistically significant difference |
off (retired) |
big |
76.6393 |
74.7915 |
+1.8478 |
[+1.3257, +2.3646] |
COMPARABLE |
NOT COMPARABLE |
historical only |
on |
big |
75.7221 |
75.5053 |
+0.2168 |
[-0.2414, +0.6713] |
COMPARABLE |
NOT COMPARABLE |
no statistically significant difference |
The live rows come from the record and the resource-matched oracle sidecar through
scripts/regenerate_arctic_paired_analysis.py; the retired rows come from the
sidecar’s preserved historical comparison. Both are regenerated from the checked-in
artifacts, so an amended measurement cannot leave this table stale.
The live oracle rows are re-decoded from the preserved upstream model bytes whenever the pinned decode resources change, so the two arms of the live comparison always consume the same dictionary and language model. The retired rows are the preserved historical comparison and are reported on their own resources; see oracle provenance.
Gap composition
This subsection is a retained earlier experiment. Its figures were computed against the 2026-08-11 oracle decode and the dictionary of that era, so they describe how the gap decomposed then; they are not the live baseline above and do not move with it.
Cross-decoding the preserved earlier-era pstrain models through this pin’s decode path shows that both model instance and decode path contribute. In the big off cell, the preserved model is +1.344 pp behind the oracle (95% CI [+0.832, +1.855]) and is -0.504 pp better than the pin-retrained model (95% CI [-0.958, -0.050]): 72.7% of the pin gap is present in the preserved model on the modern path and 27.3% is the retraining increment. In the big on cell, the corresponding values are +0.787 pp versus oracle (95% CI [+0.253, +1.320]) and -0.427 pp versus the pin model (95% CI [-0.881, +0.027]), or 64.8% and 35.2% of the pin gap respectively.
Across eras, moving the preserved pstrain model to the pin path changes WER by only +0.070 pp off and +0.157 pp on, while the byte-identical upstream oracle improves by 0.597 pp off and 0.504 pp on. The decode path is therefore model-sensitive; the differential decode-path response supplies most of the era-to-era delta movement, and pin retraining supplies a further roughly 0.43–0.50 pp.
Forward gate
Future runs compare matched pairs against the pinned per-utterance rows. For each live cell the gate computes the paired 95% bootstrap interval of the current run minus the pinned rows, in WER percentage points, and passes only when the upper end of that interval is at or below zero. That is a non-inferiority test with a zero margin: a run passes when it reproduces the pinned rows exactly, or when its interval lies entirely at or below zero so the data rule out any worsening. An interval that straddles zero fails, even when the point estimate favors the new run, because that change cannot be told apart from a regression. The gate never absorbs a change to the recorded results: moving them is always a deliberate re-pin. The live cells’ standing against the preserved upstream models is part of this documented baseline and is not itself a regression; the retired cells’ larger gap is history and is not a target.
Both arms of the live comparison must stay on the same decode resources. When
the pinned dictionary or language model changes, re-decode the preserved oracle
models with scripts/regenerate_arctic_oracle.py before regenerating the paired
analysis; make pin-check refuses arms that name different resource bytes.
Decode-path transport
Earlier program measurements used precomputed-feature batch decoding rather than this WAV decode path. Those measurements remain bound to that path: a byte-identical acoustic model shifts by 0.50–0.60 pp WER between the two paths. This pin consequently records its full decode identity. Future comparisons must either hold that identity fixed or cross-decode the compared models.
Training skips and decode coverage
The live pin has no accepted training skips. In particular, arctic_a0587
trains under product defaults and neither former exception fires. The
training.accept_arctic_a0587_known_skip knob is retained solely to describe
the retired off profile’s provenance; that profile carries the true value,
while the live on profile disables it (and the a0302 exception band). Thus
all live cells run exception-free.
A live cell must decode every utterance its denominator names. A cell’s
absolute WER is summed from its per-utterance rows, and those rows are the
decoded subset, so a live cell whose decoded fell below its
decode_denominator would publish a headline rate over a denominator its own
coverage fields contradict, and a comparison against a record carrying the
same shortfall would report no difference at all. validate_record therefore
refuses to pin such a cell, which places the refusal ahead of record emission,
fresh-record adoption, and make config-check alike. The retired off cells
are not re-gated: they were measured on a path that no longer exists, and
their coverage is a fact about the past.
The paired comparison stays ungated, because it is sound under a shortfall:
paired_delta_ci requires identical utterance sets, so lost coverage cannot
bias the delta between the arms. A run that transiently decodes fewer
utterances than the record still compares, and drift from the pinned
decoded/denominator is surfaced in field_differences beside the row for
deliberate record adoption rather than raising or failing the WER gate. Such a
run simply cannot become the pin.
Condition contract maintenance
The record owns the condition fields it contains. A changed or removed pinned value is fatal;
a live field added after the record was written is reported as uncovered but does not masquerade
as benchmark drift. make config-check runs this authentication in the ordinary suite. Adopt
newly added fields deliberately, without changing existing pins or any measured result, with
python scripts/check_arctic_pin.py --adopt-uncovered.
Fresh-record adoption also requires exact full-cell equality for both retired
off cells. Every serialized field in each retired cell—including bootstrap
summaries and any future field—is stable; any addition, removal, or value
change is refused.
Replicability
The benchmark uses canonical ordering and is deterministic. Native Mach-O fingerprints omit randomized UUID and derived code-signature bytes, and model identity covers every acoustic-model artifact while excluding timing, provenance, and completion metadata. The same checkout and resource band therefore produce the same pinned content identities; environment identity is recorded separately.