Everything runs on this machine: models download to browser cache, audio never leaves the device. Protocol: 10 words × three clips each (correct read, minimal-pair error, sounded-out) + 5 silence clips, named word_condition. Budgets: load ≤10s, ≤700ms/clip, ≥80% correct-vs-pair discrimination, silence → nothing.
| clip | model | ms | phonemes | target | dist | verdict |
|---|