> ML_LITERATURE_ATLAS_v1.0
Research Literature Atlas
253 qualified literature records from foundational statistical learning to frontier reasoning LLMs: verified DOIs, arXiv IDs, and original bilingual syntheses.
Least squares quantization in PCM
Classic signal processing paper detailing the iterative centroid relocation algorithm universally known today as k-means.
k-means++: The Advantages of Careful Seeding
Seminal algorithms paper establishing k-means++, accelerating convergence and eliminating catastrophic local minima through smart centroid initialization.
A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise
Classic KDD paper introducing DBSCAN, revolutionizing clustering by freeing algorithms from spherical assumptions and predefined k cluster counts.
Density-Based Clustering Based on Hierarchical Density Estimates
Major advance in density clustering introducing HDBSCAN, automatically extracting the most prominent clusters of varying densities.
Isolation Forest
Groundbreaking paper introducing Isolation Forest, providing linear time complexity and high resilience to swamping and masking in outlier detection.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Widely cited paper presenting UMAP, achieving superior scaling and better preservation of global data topology compared to t-SNE.
Visualizing Data using t-SNE
Landmark visualization paper introducing t-SNE, which became the standard exploratory tool for high-dimensional representations in deep learning.
Learning representations by back-propagating errors
Seminal Nature letter showing how back-propagation solves the credit assignment problem in multi-layer perceptrons, igniting connectionist AI.
Gradient-Based Learning Applied to Document Recognition
Masterpiece IEEE paper presenting LeNet-5, convolutional neural networks, and Graph Transformer Networks applied to automated check reading.
ImageNet Classification with Deep Convolutional Neural Networks
Historic NeurIPS paper presenting AlexNet, shattering the ImageNet benchmark with a 15.3% top-5 error rate and initiating modern AI.
Deep Residual Learning for Image Recognition
The most cited computer science paper of the 2010s, introducing ResNet and solving the deep network degradation problem with identity shortcut connections.
Adam: A Method for Stochastic Optimization
The landmark optimization paper introducing Adam, which computes individual adaptive learning rates for different parameters from estimates of first and second moments.
Decoupled Weight Decay Regularization
Influential optimization paper fixing the broken implementation of L2 weight decay in Adam, restoring its generalization capability to match SGD.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Breakthrough paper showing how normalizing mini-batch activations dramatically accelerates training speed and stabilizes deep convolutional networks.
Layer Normalization
Foundational normalization paper enabling sequence modeling and Transformer self-attention by decoupling normalization from mini-batch statistics.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Comprehensive JMLR paper establishing Dropout, the ubiquitous regularization technique preventing overfitting in deep neural networks.
Attention Is All You Need
The most influential AI paper of the 21st century, introducing the Transformer and multi-head dot-product self-attention to replace recurrent neural networks.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Landmark NAACL paper introducing BERT, pioneering the pre-train and fine-tune paradigm that dominated natural language processing.
