bag-of-words / one-hot vectors are V-dimensional (100K-level) and semantics-free — they cannot capture relations like king − man + woman ≈ queen; embeddings compress semantics into
d≪V dense space with cosine similarity in
[−1,1], replacing keyword matching with semantic search. With V=100K, d=4096 the embedding table alone is ~400M params (~1.6GB FP32), showing why high-dimensional sparse representations are infeasible.