Skip to content

Add AL evaluation/benchmark papers (NeurIPS 2023, TMLR 2025, TMLR 2026) - #5

Open
sten2lu wants to merge 1 commit into
SupeRuier:masterfrom
sten2lu:add-al-evaluation-papers
Open

sten2lu wants to merge 1 commit into
SupeRuier:masterfrom
sten2lu:add-al-evaluation-papers

Conversation

@sten2lu

@sten2lu sten2lu commented Mar 15, 2026

Copy link
Copy Markdown

Hi! Adding three papers on AL evaluation methodology:

  1. [NeurIPS 2023] Navigating the Pitfalls of AL Evaluation — large-scale benchmark identifying 5 common evaluation flaws. Added to: NeuraIPS.md, Benchmarks section, deep_AL_criticism.md.

  2. [TMLR 2025] nnActive — largest AL benchmark for 3D biomedical segmentation (~150k GPU hours, 8 methods × 4 datasets). Code: https://github.com/MIC-DKFZ/nnActive, Results: https://huggingface.co/nnActive. Added to: Benchmarks section, deep_AL_criticism.md.

  3. [TMLR 2026] Finally Outshining the Random Baseline — simple effective AL solution consistently beating random in 3D medical imaging. Code: https://github.com/MIC-DKFZ/nnActive, Results: https://huggingface.co/nnActive. Added to: Benchmarks section, deep_AL_criticism.md.

Happy to adjust placement or format if needed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant