I work in CRISPR discovery research. This is one of the most exaggerated nothing-burgers ever and would be laughed out of the room if a human scientist attempted to publish something like this. (cont.)
T′ᵢⱼₖ = (∂xᵃ/∂x′ᵢ)(∂xᵇ/∂x′ⱼ)(∂xᶜ/∂x′ₖ) Tₐᵦ𝑐
Woldemar Voigt gave the objects their modern name in 1898 while studying the stresses that stretch a crystal; the same transformation law later let Einstein describe how matter stretches spacetime.
A cube of numbers becomes a tensor only when its components obey this exact product of partial derivatives under every change of coordinates. A scalar needs none of them; a vector needs one Jacobian factor; a matrix needs two.
Each extra index multiplies the transformation so that the entire array continues to describe the same geometric or physical quantity, independent of the axes chosen.
"A Mathematical Explanation of Transformers" is a recent paper that develops a rigorous mathematical framework for understanding the architecture behind Transformers and large language models.
It interprets the Transformer as a discretization of a continuous integro-differential equation, with self-attention represented as a non-local integral operator, layer normalization as a projection onto a constrained set, and feedforward layers and activation functions incorporated into the same mathematical framework, then uses operator splitting and numerical discretization to recover the standard Transformer architecture and extend the formulation to multi-head attention, Vision Transformers, and convolutional Transformers.
I've already shared several resources on the mathematics behind neural networks, Transformers, and LLMs, but there always seems to be something new and interesting to explore in this area.
https://t.co/18IZwRRDZS
I’m about to give a lecture to the Columbia Math Dept, but honestly I just wanna ask THE STUDENTS how they feel about AI frontrunning their life’s work & each other. 😂 Here’s what happened in plain English:
2 groups of mathematicians made a lifetime breakthrough one of the most famous problems that had gone unsolved for almost 100 years.
But they only solved a much easier version of the problem w/lots of AI assistance, & then a false rumor went around that Anthropic had solved the problem because one of the authors was working at Anthropic lmao.
So then OpenAI decided to frontrun them & spend $20 million of compute to solve the problem. 😏
And since it’s a famous problem there’s a $1M prize for it like the Nobel.
Then the OpenAI dude told the group leader you can share the prize with us & we’ll even let you claim first authorship but you need to drop the name of your Anthropic coauthor from the paper.
Then the group leader got mad because that’s super unethical & felt like they secretly used his ChatGPT data to figure out what he was working on to frontrun him. At least that’s the Anthropic perspective.
The problem is so hard that the original two authors would’ve gotten a top math prize & generational prestige for just solving the easier version.
But OpenAI just brute forced it while seemingly stealing their idea and being unethical and everything.
If not for the drama the scientific achievement would’ve been comparable to curing a form of cancer in their field‼️
But instead no one’s questioning that it got solved, & instead it devolved into this shitshow. 😌
Now a lot of mathematicians are saying the hardest problems that they spend a lifetime working on are gonna get solved in a year or two. 🙃
And some mathematicians are now saying it’s like doping and they should ban using AI on some types of problems to “maintain the challenge.”
@rynorhn I can't speak for them obviously but in academia it is almost always official policy to never put anything confidential or new into LLMs unless they are institutionally made because WE KNOW the corporations HAVE, CAN AND WILL steal intellectual property / confidential research.
@rynorhn Grigori Perelman: “I can’t say I’m outraged. Other people do worse. Of course, there are many mathematicians who are more or less honest. But almost all of them are conformists. They are more or less honest, but they tolerate those who are not honest.”
Don’t tolerate it.
One second-order equation determines the straightest curves a manifold will admit:
d²xᵏ/dt² + Γᵏᵢⱼ(x) (dxⁱ/dt)(dxʲ/dt) = 0.
Riemann supplied the metric that generates those symbols in 1854; Christoffel isolated the connection coefficients fifteen years later.
The extra terms, built from the Christoffel symbols of the metric, cancel the fictitious accelerations that the coordinates themselves produce, so the left-hand side is the covariant acceleration and vanishes.
The resulting curves are the geodesics - great circles on the sphere, meridians, the world-lines of freely falling particles.