Datasets & Benchmarks
Source note: this page is educational and was last checked against official portal landing pages on 2026-06-04. Dataset sizes, releases, access levels, and schemas change; verify the source portal before quoting numbers or building a benchmark.
Cancer datasets are not interchangeable. A useful dataset entry should say what was measured, how samples were selected, whether data are open or controlled-access, which version was used, and what the dataset is not suitable for.
High-value starting points
| Resource | Use it for | Cautions |
|---|---|---|
| NCI Genomic Data Commons / TCGA | Cohort discovery, harmonized genomic files, clinical metadata, TCGA/TARGET-style learning | Counts and file availability change by project and release; controlled data require authorization |
| cBioPortal / MSK-IMPACT studies | Exploratory clinico-genomic queries and mutation/copy-number summaries | Study versions and sample filters matter; not every cohort is a benchmarking truth set |
| AACR Project GENIE | Real-world clinico-genomic registry data | Use the specific public release and data guide; clinical completeness varies by contributing center |
| TCIA | Cancer imaging collections, radiology/pathology imaging research | Imaging protocols, segmentations, labels, and linked clinical data vary by collection |
| Human Cell Atlas | Single-cell and spatial reference context | Not oncology-specific by default; use tissue, donor, assay, and disease metadata carefully |
Sources: [1], [2], [3], [4], [5]
What to record before using a dataset
- Source URL, portal release, download date, and accession/study ID.
- Open vs controlled-access status and data-use restrictions.
- Sample inclusion/exclusion rules and duplicate handling.
- Assay type, platform, reference genome, pipeline, and normalization.
- Endpoint definitions, censoring rules, and missing-data handling.
- Whether labels are diagnostic, prognostic, predictive, synthetic, weakly supervised, or manually curated.
Benchmark rules
A dataset becomes a benchmark only after the task is locked:
| Task | Minimum benchmark definition |
|---|---|
| Variant calling | Reference genome, truth set, region mask, variant classes, caller versions, scoring metric |
| Expression analysis | Count matrix provenance, batch design, normalization, contrast, multiple-testing policy |
| Survival modeling | Start time, event, censoring, follow-up, train/test split, calibration metric |
| Imaging AI | DICOM metadata, segmentation source, preprocessing, scanner/site split, external validation |
| Trial matching | Source trial registry date, patient fact schema, criterion-level labels, human review policy |
Do not treat a large public dataset as automatically fair, representative, or clinically validated. Public availability is not the same as benchmark readiness.
Local sample data
The file data/samples/sample_cancer_data.csv is synthetic demonstration data. It is not an incidence cohort, survival cohort, treatment-response cohort, or clinical benchmark. See data/samples/README.md in the repository before using it in examples.
See also
References
- NCI Genomic Data Commons. GDC Data Portal. https://portal.gdc.cancer.gov/
- National Cancer Institute. The Cancer Genome Atlas Program (TCGA). https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga
- AACR Project GENIE. AACR Project GENIE Data. https://www.aacr.org/professionals/research/aacr-project-genie/aacr-project-genie-data/
- The Cancer Imaging Archive. TCIA collections. https://www.cancerimagingarchive.net/
- Human Cell Atlas. HCA Data Portal. https://data.humancellatlas.org/