> ML_VERI_KUMELERI_v1.0
Veri Kümeleri Atlası
Yapay zeka modellerinin eğitildiği ve değerlendirildiği 103 doğrulanmış veri kümesi: lisans katmanları, veri gizliliği riskleri, önyargı ve sızıntı analizleri.
SWE-bench: Real-World GitHub Software Engineering Issues
Software Engineering & Autonomous Agents · 2,294 real-world GitHub issues and pull request test patches across 12 popular open-source Python repos · Princeton NLP (Jimenez et al.)
Natural Questions (NQ - Google Research 2019)
Open-Domain Question Answering & Search · 307,373 training examples from real Google Search queries paired with Wikipedia pages · Google Research (Kwiatkowski et al.)
HumanEval: Handcrafted Python Coding Benchmark
Code Generation & Functional Correctness · 164 handwritten Python programming problems with docstrings, solutions, and unit test cases · OpenAI (Chen et al. 2021)
MBPP: Mostly Basic Python Problems (Austin et al. 2021)
Introductory Code Synthesis · 974 entry-level Python programming problems crowd-sourced with automated test cases · Google Research
TruthfulQA: Measuring Model Imitation of Falsehoods
Epistemic Truthfulness & Alignment · 817 adversarial questions across 38 categories (health, law, conspiracies, superstitions) · Oxford / Evans Lab (Lin et al. 2022)
OpenAssistant Conversations (OASST1)
Instruction Tuning & Conversational RLHF · 161,443 messages in 66,497 conversation trees across 35 languages (13,500+ volunteer contributors) · LAION / Large-scale AI Open Network (Köpf et al. 2023)
LibriSpeech ASR Corpus (Panayotov et al. 2015)
Speech Recognition & Acoustic Modeling · 1,000 hours of 16kHz read English audiobooks with aligned text transcripts · Johns Hopkins University / Daniel Povey
Mozilla Common Voice (Multilingual Speech Corpus)
Multilingual Speech Recognition · 30,000+ hours across 120+ languages recorded by global crowdsourced volunteers · Mozilla Foundation
AudioSet: A Large-Scale Dataset of Manually Annotated Audio Events
General Audio & Sound Event Detection · 2,084,320 10-second YouTube sound video clips spanning 527 sound classes organized in an ontology · Google Research (Gemmeke et al. 2017)
LAION-5B: Open Large-Scale Multimodal Dataset
Multimodal Foundation Pretraining · 5.85 billion CLIP-filtered image-URL and alt-text pairs (2.3B English, 2.2B multilingual) · LAION e.V. (Schuhmann et al. 2022)
MS COCO Captions (Chen et al. 2015)
Vision-Language & Image Captioning · 330,000 images with 5 independent human crowdsourced English captions per image (1.5M+ captions) · Microsoft / COCO Consortium
Visual Genome: Connecting Language and Vision (Krishna et al. 2017)
Visual Relational Knowledge & Scene Graphs · 108,077 images with 5.4M region descriptions, 2.3M visual relationships, and 2.8M attributes · Stanford University (Ranjay Krishna et al.)
MIMIC-IV (Medical Information Mart for Intensive Care IV)
Healthcare & Clinical Informatics · Hospital stays for >300,000 patients admitted between 2008-2019 · MIT PhysioNet / Beth Israel Deaconess Medical Center
CheXpert: A Large Chest Radiograph Dataset (Irvin et al. 2019)
Medical Imaging & Radiology · 224,316 chest radiographs of 65,240 patients with radiologist report extractions · Stanford ML Group (Stanford Hospital)
Cora Citation Network Benchmark
Graph Machine Learning & Bibliometrics · 2,708 scientific publications and 5,429 citation links across 7 machine learning sub-fields · McCallum et al. / Automating the Construction of Internet Portals
OGB: ogbn-arxiv Citation Network (Hu et al. 2020)
Large-Scale Graph Machine Learning · 169,343 arXiv computer science papers and 1,166,243 directed citation edges across 40 categories · Open Graph Benchmark / Stanford University (Jure Leskovec et al.)
OGB: ogbg-molpcba (PubChem BioAssay Molecular Graphs)
Computational Chemistry & Drug Discovery · 437,929 small molecules represented as 2D molecular graphs with 128 multi-task biological targets · Open Graph Benchmark / Stanford University / MIT
Electricity Load Diagrams (ECL 2011-2014)
Energy & Utility Demand Forecasting · 140,256 15-minute electricity consumption records across 321 Portuguese clients · UCI Machine Learning Repository / Artur Trindade
