> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Regression Shrinkage and Selection via the Lasso
Landmark statistical paper creating the Lasso, solving multicollinearity and high-dimensional feature selection via L1 penalization.
Ridge Regression: Biased Estimation for Nonorthogonal Problems
Foundational paper introducing Ridge regression to stabilize parameter estimates in ill-conditioned and multicollinear regression settings.
Regularization and variable selection via the elastic net
Seminal regularized regression paper showing that combining L1 and L2 penalties yields stable feature selection on strongly correlated predictors.
C4.5: Programs for Machine Learning
Classic machine learning book establishing the C4.5 decision tree algorithm, universally voted one of the top 10 algorithms in data mining.
A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting (AdaBoost)
Gödel Prize-winning computer science paper introducing AdaBoost, founding the field of ensemble boosting by adaptively re-weighting misclassified points.
Extremely Randomized Trees (Extra-Trees)
Key ensemble learning paper presenting Extra-Trees, achieving competitive tabular accuracy with minimal computational overhead.
LIBLINEAR: A Library for Large Linear Classification
Essential machine learning systems paper detailing LIBLINEAR, the engine underpinning linear models in scikit-learn and high-scale production systems.
LIBSVM: A Library for Support Vector Machines
ACM TIST classic paper introducing LIBSVM, the foundational C++ kernel SVM software used worldwide across all major programming languages.
Scikit-learn: Machine Learning in Python
Monumental JMLR paper establishing scikit-learn, the universal standard library for classical machine learning and data science globally.
No Free Lunch Theorems for Optimization
Profound foundational theory proving that no universal optimal algorithm exists; inductive bias tailored to domain structure is mandatory.
Nonlinear Component Analysis as a Kernel Eigenvalue Problem (Kernel PCA)
Foundational non-linear dimension reduction paper showing that the kernel trick generalizes linear eigenvalue problems to arbitrary Hilbert feature spaces.
A Stochastic Approximation Method
The mathematical origin of all stochastic gradient descent and mini-batch machine learning optimization.
Understanding the difficulty of training deep feedforward neural networks (Xavier / Glorot Initialization)
Seminal AISTATS paper resolving vanishing/exploding activations in deep networks by mathematically preserving variance across forward and backward passes.
Focal Loss for Dense Object Detection (RetinaNet)
Marr Prize-winning computer vision paper introducing Focal Loss, enabling single-stage object detectors (RetinaNet) to surpass two-stage detectors.
Gaussian Error Linear Units (GELUs)
Essential activation paper establishing GELU, the standard non-linear activation used in BERT, GPT-2, GPT-3, ViT, and RoBERTa.
RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE)
The defining positional encoding technique adopted in LLaMA, Mistral, Qwen, DeepSeek, and modern foundation models for seamless context length scaling.
U-Net: Convolutional Networks for Biomedical Image Segmentation
The most influential biomedical vision paper of all time, whose U-Net architecture is both the medical segmentation standard and the core backbone of DDPM diffusion.
Densely Connected Convolutional Networks (DenseNet)
CVPR Best Paper Award winning work introducing DenseNet, maximizing feature reuse through dense iterative channel concatenations.
