Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.
Evaluation Protocol · Training & Decoding
HearInContext evaluates how speech recognition models use implicit semantic cues from dialogue. Shared-audio comparisons distinguish contextual disambiguation from explicit word hints.
- Mandarin and English, with no-context, implicit, unrelated, and explicit conditions.
- Unified CER/WER and target recall, with breakdowns by condition and speaker.
- Qwen3-ASR fine-tuning and decoding entrypoints, plus paired bootstrap tools.
- CPU-only scoring; GPU-based model training and decoding.
Run from the repository root using Python 3.12:
conda create -n hearincontext python=3.12 -y
conda activate hearincontext
python -m pip install -r requirements.txtFor the training and decoding environment, see the Qwen guide.
Training and test data are distributed separately on Hugging Face. The dataset card covers data construction, speech synthesis, reference speakers, and data licensing.
Prepare a test manifest, test.jsonl, and model predictions, predictions.jsonl. Each prediction must retain the sample's full example_id:
{"example_id":"demo::C1::voice1","hypothesis":"Please check the flour."}Run scoring:
python evaluation/evaluate.py \
--manifest test.jsonl \
--hypotheses predictions.jsonl \
--output outputs/metrics.jsonIDs must be unique in each file and match exactly across files. Empty transcripts are scored normally. Use a new output path.
Results are grouped by language and condition. error_rate and target_recall are proportions; multiply by 100 for percentages. Add --observations outputs/rows.jsonl to save per-example statistics.
Train from an official Qwen3-ASR base model, then decode with the selected checkpoint:
We provide code and data, but do not distribute fine-tuned weights.
Run tests:
python -m unittest discover -s tests@misc{gao2026hearincontext,
title = {HearInContext: A Benchmark for Implicit Context in Speech Recognition},
author = {Gao, Yifan and Tian, Yao and Suo, Hongbin},
year = {2026}
}- Qwen3-ASR for the model, training, and inference implementations.
- Whisper / Transformers for English text normalization.
The code is licensed under Apache-2.0. Third-party code retains its original copyright and license notices; see NOTICE.
