⚠️ IMPORTANT: ALL DATA IS SYNTHETICThis dataset contains entirely synthetic medical notes generated using large language models. No real patient data was used. These notes are created for research and educational purposes only. While grounded in clinical guidelines and reviewed by a qualified GP, they may contain inaccuracies and should not be used for clinical decision-making.
SynGP500 is a clinician-curated collection of 500 synthetic Australian general practice medical notes created to support machine learning and natural language processing research in primary care.
The Problem: Access to clinical text data is a significant challenge for healthcare NLP research due to privacy regulations and ethical constraints. Publicly available datasets for Australian general practice are particularly limited.
The Solution: SynGP500 addresses this challenge through systematic multi-dimensional grounding, combining curriculum-based clinical breadth, epidemiologically-calibrated prevalence, contextually diverse encounter settings, and authentic documentation complexity.
The Result: A dataset exhibiting realistic linguistic variation alongside the clinical complexity and contextual constraints of genuine general practice, providing researchers and educators with a privacy-preserving resource for developing and evaluating clinical NLP methods.
📄 Full paper on arXiv: SynGP500
- 500 synthetic consultation notes (average 606 ± 257 words, range 213–1,444)
- Adult and elderly focus (18+ years) - pediatric cases not included in this pre-release version
- Clinical presentations: Aligned with RACGP registrar curriculum and BEACH study prevalence patterns
- Nine consultation settings: Standard clinic, bulk billing, RACF, telehealth, home visits, after-hours, community health, Aboriginal health services, mobile outreach
- Geographic diversity: Metropolitan, regional, rural, and remote (MM1-MM7 remoteness classifications)
- Documentation styles: Multiple synthetic clinician personas with varying patterns
- Epidemiologically validated: Case distribution closely matches BEACH study data (within ±1-2% for most presenting complaint categories)
- Demonstrated realism: Natural typo rate (0.83%), high stylometric diversity (MATTR 0.858–0.946), realistic length variation (CV 0.42-0.47)
- Semantic diversity validated: Note-level embeddings show broad distribution (cosine similarity mean 0.52, range 0.09-0.95), providing evidence against mode collapse typical of naïve LLM generation
- Medical authenticity: 48.3% medical term density, SNOMED-CT-AU coded for systematic ontological coverage
- NER validated: Grouped Type F1 score of 0.6951 (+14.7% improvement over baseline) using MedCAT on GP-authored hypothetical test cases
To access the dataset:
# Clone the repository
git clone https://github.com/pisong314/syngp500.git
cd syngp500
# Browse the notes
ls notes/
# View a sample note
cat notes/14669001_0093_Acute_kidney_injury.txtAll 500 synthetic medical notes are located in the /notes directory as plain text files (UTF-8 encoded). Each filename follows the pattern: {SNOMED_code}_{ID}_{condition_name}.txt
- Realistic consult complexity - See actual examples how complexities of real consults are reflected in this dataset.
- BEACH Epidemiological Comparison - Case distribution validation against real Australian GP data
- Stylometric Analysis - Linguistic diversity metrics (MATTR, typo rates, style variation)
- Semantic Diversity Analysis - Embedding space analysis showing broad distribution and absence of mode collapse
- NER Performance - MedCAT evaluation results (Grouped Type F1: 0.6951, +14.7% improvement)
- Generation Architecture - How the synthetic notes were created, including LLM-based generation, clinical grounding, and quality assurance
- Scalability - Framework scalability
- Use Cases - Recommended applications and precautions
- Limitations - Important constraints and scope
- Contributing - How to report issues and improve the dataset
Before using this dataset:
- Synthetic data only - No real patient data. Created entirely using LLMs and clinical knowledge.
- Limited validation - Single-clinician review. May contain clinical inaccuracies despite careful curation.
- Not RACGP endorsed - Independent research project, not affiliated with RACGP.
- Primary use: ML/NLP research - For training models, benchmarking, proof-of-concept development.
- Educational use requires review - Educators must verify clinical accuracy of each case before teaching.
- Report issues - Found an error? Email piyawoot.song@gmail.com to help improve the dataset.
This dataset is released under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).
- ✅ Free for research and education (non-commercial use)
- ✅ Share and adapt with appropriate attribution
- ✅ Must use the same license for derivative works
- 📧 Contact for commercial use or larger synthetic datasets
For full license details, see the LICENSE file or visit: https://creativecommons.org/licenses/by-nc-sa/4.0/
If you use this dataset, please cite the arXiv paper. You can use GitHub's "Cite this repository" button or the BibTeX below:
@misc{songsiritat2025syngp500,
author = {Songsiritat, Piyawoot},
title = {SynGP500: A Clinically-Grounded Synthetic Dataset of Australian General Practice Medical Notes},
year = {2025},
note = {arXiv preprint arXiv:2512.15259},
eprint = {2512.15259},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2512.15259}
}Dr Piyawoot Songsiritat
MBBS, FRACGP
Clinical NLP Researcher
📧 Email: piyawoot.song@gmail.com
For inquiries about:
- Collaboration opportunities
- Larger synthetic datasets
- Technical questions about dataset generation
The author acknowledges Aboriginal and Torres Strait Islander peoples as the Traditional Custodians of Australia and pays respect to Elders past, present, and emerging. This dataset includes synthetic cases that represent the diversity of patients seen in Australian general practice, including Aboriginal and Torres Strait Islander peoples. While these are entirely synthetic cases, they reflect the importance of culturally safe healthcare delivery and recognition of health disparities affecting Indigenous Australians.
This dataset was created to support the clinical NLP research community and medical education. The author acknowledges the importance of evidence-based guidelines from RACGP, Therapeutic Guidelines, and other Australian clinical resources that informed the synthetic note generation process.
-
Australian Digital Health Agency. SNOMED CT-AU (Australian extension of SNOMED CT). Available from: https://www.healthterminologies.gov.au
-
Britt H. BEACH--bettering the evaluation and care of health: a continuous national study of general practice activity. Commun Dis Intell Q Rep. 2003;27(3):391-393. doi:10.33321/cdi.2003.27.68
-
Kraljevic Z, Searle T, Shek A, Roguski L, Noor K, Bean D, Mascio A, Zhu L, Folarin AA, Roberts A, Bendayan R, Richardson MP, Stewart R, Shah AD, Wong WK, Ibrahim Z, Teo JT, Dobson RJB. Multi-domain clinical natural language processing with MedCAT: The Medical Concept Annotation Toolkit. Artif Intell Med. 2021;117:102083. doi:10.1016/j.artmed.2021.102083
-
Rajotte JF, Bergen R, Buckeridge DL, El Emam K, Ng R, Strome E. Synthetic data as an enabler for machine learning applications in medicine. iScience. 2022;25(11):105331. doi:10.1016/j.isci.2022.105331
-
Royal Australian College of General Practitioners. 2022 RACGP curriculum and syllabus for Australian general practice (6th ed.). 2022. Available from: https://www.racgp.org.au/education/education-providers/curriculum/curriculum-and-syllabus/home
- Semantic diversity validation added (UMAP embeddings)
- Framework scalability discussion added
- NER validation added: MedCAT evaluation on clinician-authored fictional notes with clinician annotations
- More organised README.md
- Initial release: 500 synthetic GP notes
Last Updated: December 2025