The SCUT-HCCDoc Dataset for the research of offline handwritten Chinese text recognition (HCTR) in camera-captured document images is now released by Deep Leaning and Visual Computing Lab of South China University of Technology. The dataset can be downloaded through the following link:
- Baidu Cloud (We will give the download link later)
- OneDrive
Note: The SCUT-HCCDoc dataset can only be used for non-commercial research purpose. For scholars or organization who wants to use the SCUT-EPT database, please first fill in this Application Form and send it via email to us (eelwjin@scut.edu.cn). When submiting the application form to us, please list or attached 1-2 of your publications in recent 6 years to indicate that you (or your team) do research in the related research fields of OCR, handwriting analysis and recognition, document image processing, and so on. We will give you the decompression password after your letter has been received and approved.
The SCUT-HCCDoc Dataset contains 12,253 camera-captured natural images with 116,629 text lines and 1,155,801 characters. According to different application scenes, SCUT-HCCDoc can be roughly divided into five subsets:
- HCCDoc-WT: images of traditional Chinese characters;
- HCCDoc-WS: images of simplified Chinese characters without a formatted background;
- HCCDoc-WSF: images of simplified Chinese characters with the formatted background;
- HCCDoc-SN: images of student notes;
- HCCDoc-EP: images of examination papers.
The sample distribution of five subsets is shown below:
The comparison of the five subsets of SCUT-HCCDoc in terms of the character number and text line box number (ABN is average box number; ACN is average character number) is shown below.
The following are some page/text level images of SCUT-HCCDoc:
The diversity of SCUT-HCCDoc can be described in three levels:
- Image-level diversity: image appearance and geometric variances caused by camera-captured settings (such as perspective, background, and resolution) and different applications (such as note-taking, test papers, and homework);
- Text-level diversity: variances of text line length, rotation, etc.;
- Character-level diversity: variances of character categories (up to 6,109 classes with additional English letters, and digits), character size, individual writing style, etc.
For example, the following image shows the number of character instances for the 50 most frequently observed character categories in the SCUT-HCCDoc.
Please consider to cite our paper when you use our dataset:
@article{zhang2020hccdoc,
author = {Zhang, Hesuo and Liang, Lingyu and Jin, Lianwen},
title = {SCUT-HCCDoc: A New Benchmark Dataset of Handwritten Chinese Text in Unconstrained Camera-captured Documents},
journal = {Pattern Recognition},
year = {2020},
publisher = {Elsevier}
}
For any quetions about the dataset please contact the authors by sending email to Hesuo Zhang (eehesuo.zhang@mail.scut.edu.cn), Lingyu Liang (lianglysky@gmail.com) or Prof. Jin (eelwjin@scut.edu.cn).







