> ML_DATASET // OPENML-CC18-BENCHMARK-SUITE_v1.0
OpenML-CC18 Curated Classification Benchmark Suite
OpenML Consortium (Bischl et al.) · Meta-Learning & AutoML · 72 diverse tabular datasets across scientific, industrial, and social domains
Meta-Learning & AutoMLCC-BY-4.072 diverse tabular datasets across scientific, industrial, and social domainsopen
Dataset Profile & Characteristics
Label Type:Binary and multiclass classification targets
Languages:en
License Tier:permissive-open-source
Modalities:tabular
Intended Use
- AutoML benchmarking and rigorous comparison of tabular classifiers (XGBoost vs LightGBM vs CatBoost vs Deep Learning)
Prohibited / Discouraged Use
- Evaluating computer vision or unstructured NLP backbones
Bias, Leakage & Privacy Risk Analysis
Privacy / Sensitive Data Risks:
Curated open-access scientific datasets; all personal datasets are fully anonymized.
Known Bias:
Aggregated collection overweights mid-sized benchmark tabular tasks.
Known Benchmark Leakage:
Strict 10-fold cross-validation splits standardized via OpenML task definitions.
