Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

English | 中文

HearInContext

A Benchmark for Implicit Context in Speech Recognition

GitHub arXiv Hugging Face Dataset License: Apache-2.0

HearInContext overview

Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.

Evaluation Protocol · Training & Decoding

HearInContext evaluates how speech recognition models use implicit semantic cues from dialogue. Shared-audio comparisons distinguish contextual disambiguation from explicit word hints.

  • Mandarin and English, with no-context, implicit, unrelated, and explicit conditions.
  • Unified CER/WER and target recall, with breakdowns by condition and speaker.
  • Qwen3-ASR fine-tuning and decoding entrypoints, plus paired bootstrap tools.
  • CPU-only scoring; GPU-based model training and decoding.

Installation

Run from the repository root using Python 3.12:

conda create -n hearincontext python=3.12 -y
conda activate hearincontext
python -m pip install -r requirements.txt

For the training and decoding environment, see the Qwen guide.

Dataset

Training and test data are distributed separately on Hugging Face. The dataset card covers data construction, speech synthesis, reference speakers, and data licensing.

Quick Evaluation

Prepare a test manifest, test.jsonl, and model predictions, predictions.jsonl. Each prediction must retain the sample's full example_id:

{"example_id":"demo::C1::voice1","hypothesis":"Please check the flour."}

Run scoring:

python evaluation/evaluate.py \
  --manifest test.jsonl \
  --hypotheses predictions.jsonl \
  --output outputs/metrics.json

IDs must be unique in each file and match exactly across files. Empty transcripts are scored normally. Use a new output path.

Results are grouped by language and condition. error_rate and target_recall are proportions; multiply by 100 for percentages. Add --observations outputs/rows.jsonl to save per-example statistics.

Training & Decoding

Train from an official Qwen3-ASR base model, then decode with the selected checkpoint:

We provide code and data, but do not distribute fine-tuned weights.

Further Usage

Run tests:

python -m unittest discover -s tests

Citation

@misc{gao2026hearincontext,
  title  = {HearInContext: A Benchmark for Implicit Context in Speech Recognition},
  author = {Gao, Yifan and Tian, Yao and Suo, Hongbin},
  year   = {2026}
}

Acknowledgements

License

The code is licensed under Apache-2.0. Third-party code retains its original copyright and license notices; see NOTICE.

About

HearInContext: A Benchmark for Implicit Context in Speech Recognition

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages