Pratham Patel

Posts tagged "deep-learning"

9 posts found

X-Token: Cross-Tokenizer Knowledge Distillation from Scratch

Building cross-tokenizer distillation from the ground up — why per-position KL breaks across tokenizers, DP span alignment, the chain-rule chunk merge, the projection matrix W, and two complementary losses (P-KL and H-KL) — with a full proof of GOLD's suppressive gradient and one running 2+3 example derived step by step.

knowledge-distillationtokenizersllm

Mathematical Prerequisites for Mamba

Building the math foundations you need for Mamba — first-order linear ODEs, the integrating factor method, the matrix exponential, the integral identity used in zero-order hold, the ZOH discretization itself, the unrolled time-varying recurrence, and the 1-semiseparable matrix view — all derived step by step with one consistent leaky-integrator example.

machine-learningmambastate-space-models

From Soft Alignment to Queries, Keys, and Values: Deriving the Transformer's Attention

The Q/K/V abstraction derived from first principles — why Bahdanau's feedforward alignment collapses to a dot product, how the Transformer formalizes queries, keys, and values as separate projections, why we scale by √d_k (with a complete variance proof), what multi-head attention adds, and a full 3-token self-attention numerical walkthrough

deep-learningattentiontransformers

What Attention is Really Doing: Weighted Memory Retrieval from Scratch

Building attention from first principles — the fixed-length bottleneck that broke RNN encoder–decoders on long sentences, the weighted-sum solution introduced by Bahdanau, Cho, and Bengio (2015), alignment score derivation, softmax normalization, and a complete 3-word numerical walkthrough — all derived step by step

deep-learningattentiontransformers

Mathematical Prerequisites for the Attention Series

Building the math foundations for the attention series — the tanh function, dot products as similarity measures, the standard normal distribution, variance, independence of random variables, and why the variance of a dot product equals the vector dimension — all derived step by step with one consistent 2-dimensional example.

attentiontransformersmathematics

Mixture of Experts from Scratch — Part 3: Why MoEs Work and the Modern Landscape (2022–2024)

Why experts specialize instead of collapsing, the role of nonlinearity, exploration and router learning, load balancing theory, and a complete taxonomy of modern MoE in LLMs — all derived from first principles with concrete examples.

machine-learningmixture-of-expertsdeep-learning

Attention Residuals: Replacing Fixed Skip Connections with Learned Depth-Wise Attention

Building Attention Residuals from scratch — why standard residuals dilute information, how softmax attention over depth fixes it, the block variant that makes it practical, and the structured-matrix view that unifies everything — all derived step by step with a 4-layer running example

deep-learningtransformersarchitecture

Manifold-Constrained Hyper-Connections: Stabilizing Deep Networks Beyond ResNets (with the actual math)

From residual identity paths to Hyper-Connections and mHC — now with the paper's exact equations, fully unrolled products, and concrete numeric examples

deep-learningneural-networkslinear-algebra

Manifold-Constrained Hyper-Connections: Stabilizing Deep Networks Beyond ResNets

A deep dive into why residual connections work, how Hyper-Connections generalize them, and why constraining learned skip paths to doubly stochastic matrices solves the instability problem

deep-learningneural-networkslinear-algebra