A size distribution analysis of the pancreatic_lesion.nii.gz annotation files identifies 9 cases with ground truth lesions smaller than 50 voxels. The smallest observed lesion is 4 voxels, which at any plausible voxel spacing for pancreatic CT has no clinical correlate as a detectable lesion. These cases likely reflect annotation inconsistencies rather than true micro-lesions.
Note that this analysis covers only pancreatic_lesion.nii.gz and has not been extended to other label sources in the dataset; similar patterns may or may not exist elsewhere.
Affected Cases
The following cases were identified, sorted by lesion size ascending: PanTS_00005043 (4 voxels), PanTS_00006927 (7 voxels), PanTS_00003059 (12 voxels), PanTS_00000044 (16 voxels), PanTS_00007933 (16 voxels), PanTS_00005693 (20 voxels), PanTS_00007947 (32 voxels), PanTS_00003548 (36 voxels), and PanTS_00001819 (45 voxels).
Very small ground truth masks disproportionately inflate Dice score variance and can produce near-zero DSC for any imperfect prediction, which may distort evaluation results.

A size distribution analysis of the pancreatic_lesion.nii.gz annotation files identifies 9 cases with ground truth lesions smaller than 50 voxels. The smallest observed lesion is 4 voxels, which at any plausible voxel spacing for pancreatic CT has no clinical correlate as a detectable lesion. These cases likely reflect annotation inconsistencies rather than true micro-lesions.
Note that this analysis covers only pancreatic_lesion.nii.gz and has not been extended to other label sources in the dataset; similar patterns may or may not exist elsewhere.
Affected Cases
The following cases were identified, sorted by lesion size ascending: PanTS_00005043 (4 voxels), PanTS_00006927 (7 voxels), PanTS_00003059 (12 voxels), PanTS_00000044 (16 voxels), PanTS_00007933 (16 voxels), PanTS_00005693 (20 voxels), PanTS_00007947 (32 voxels), PanTS_00003548 (36 voxels), and PanTS_00001819 (45 voxels).
Very small ground truth masks disproportionately inflate Dice score variance and can produce near-zero DSC for any imperfect prediction, which may distort evaluation results.