The identity
⎡1 0 0⎤
⎢0 1 0⎥
⎣0 0 1⎦
does nothing, and that is its power. It is the multiplicative unit Cayley introduced when, in 1858, he turned a rectangular array of numbers into an algebra with its own addition and multiplication.
Two millennia earlier the same arrays already solved simultaneous equations by column operations in The Nine Chapters.
Seki Takakazu (1683) and Leibniz (1693) independently extracted from the square case a single number—the determinant—that decides whether a unique solution exists.
Sylvester merely named the array a matrix in 1850;
Cayley supplied the operations that made the name into structure.
The same objects now encode rotations, quantum evolution, and the weights of every neural network.
The Mathematics of Large Language Models — A Readable Guide to LLMs, Transformers, Diffusion, Neural Networks, and Generative AI: https://t.co/3sDIaIroX8
General relativity replaces gravity as a force with the geometry of spacetime. Matter and energy alter that geometry, while freely falling objects and light follow its natural paths, called geodesics.
Einstein’s field equation, Gμν + Λgμν = (8πG/c⁴)Tμν,
Most of a giant LLM can stay frozen. Only a tiny Δθ is trained.
LoRA injects low-rank A and B. QLoRA first quantizes to 4-bit. Prefix tuning learns a short token prefix. Adapters insert tiny modules between layers. Selective methods update only the weights that matter.
The regularized objective that keeps new skills from erasing old knowledge:
θ* = arg min_θ ℒ(θ; 𝒟_task) + λ‖θ − θ₀‖²
Same performance, far fewer parameters, far lower cost.
Which PEFT method have you actually used in production?