Skip to main content

> ML_DATASET // OPENML-CC18-BENCHMARK-SUITE_v1.0

OpenML-CC18 Curated Classification Benchmark Suite

OpenML Consortium (Bischl et al.) · Meta-Learning & AutoML · 72 diverse tabular datasets across scientific, industrial, and social domains

Meta-Learning & AutoMLCC-BY-4.072 diverse tabular datasets across scientific, industrial, and social domainsopen

Dataset Profile & Characteristics

Label Type:Binary and multiclass classification targets
Languages:en
License Tier:permissive-open-source
Modalities:tabular

Intended Use

  • AutoML benchmarking and rigorous comparison of tabular classifiers (XGBoost vs LightGBM vs CatBoost vs Deep Learning)

Prohibited / Discouraged Use

  • Evaluating computer vision or unstructured NLP backbones

Bias, Leakage & Privacy Risk Analysis

Privacy / Sensitive Data Risks:

Curated open-access scientific datasets; all personal datasets are fully anonymized.

Known Bias:

Aggregated collection overweights mid-sized benchmark tabular tasks.

Known Benchmark Leakage:

Strict 10-fold cross-validation splits standardized via OpenML task definitions.

Compatible Tools & Libraries