Skip to main content

> ML_ALGORITHM // CONTINUOUS-BAG-OF-WORDS-SKIPGRAM-WORD2VEC_v1.0

Word2Vec (CBOW & Skip-Gram)

Pioneering neural word embedding technique that maps vocabulary tokens into dense vector spaces capturing linear semantic and syntactic regularities.

Static Word Embeddingsself-supervisedhigh-intrinsiclarge (>100k)
Back to All Algorithms
Computational Complexity
Training Complexity:O(epochs * corpus_size * window * negative_samples)
Inference Complexity:O(1) dictionary vector lookup
Hardware Profile
CPU Friendly:Yes
Requires GPU:No
Memory Footprint:low
Interpretability & Data
Interpretability Tier:high-intrinsic
Training Data Needs:large (>100k)

Interpretability Assessment

Embedding vector geometry satisfies linear semantic analogies (e.g., King - Man + Woman = Queen).

Suitable Tasks & Supported Modalities

Suitable Tasks:
feature extractiontoken embedding
Supported Modalities:
text

Implementing Libraries

GensimRaRe Technologies / Radim Řehůřek · v4.3.3
View Spec
fastTextMeta AI Research (FAIR) · v0.9.2
View Spec

Foundational Literature

Common Pitfalls & Warnings
  • Conflating multiple meanings of polysemous words (e.g., "bank" of a river vs financial "bank") into a single static vector
  • Zero vectors or crash on out-of-vocabulary (OOV) tokens