We are auditing the public MMLU auxiliary training archive under a strict source-disjoint policy. Exact bridges to several upstream datasets can be reconstructed for parts of the archive, but the official repository and paper do not appear to provide a complete per-row source sidecar for all auxiliary_train members and aliases. Dataset/member filename alone is not a sufficient independent source unit for our protocol.
Is there an immutable public mapping from every auxiliary row to its upstream dataset member, document/article/exam form, and any known aliases? If available, could you provide the exact release/commit, field semantics, license, file size and SHA-256, preserving duplicates and one-to-many origin ambiguity? We do not need row text or answers and will not infer missing origins from content similarity.
We are auditing the public MMLU auxiliary training archive under a strict source-disjoint policy. Exact bridges to several upstream datasets can be reconstructed for parts of the archive, but the official repository and paper do not appear to provide a complete per-row source sidecar for all auxiliary_train members and aliases. Dataset/member filename alone is not a sufficient independent source unit for our protocol.
Is there an immutable public mapping from every auxiliary row to its upstream dataset member, document/article/exam form, and any known aliases? If available, could you provide the exact release/commit, field semantics, license, file size and SHA-256, preserving duplicates and one-to-many origin ambiguity? We do not need row text or answers and will not infer missing origins from content similarity.