Image foundation models never stop surprising!😮
And this time, it's not the usual big tech players.
A few days ago, @robbyant_brain released LingBot-Vision, a new family of vision encoders built natively for dense spatial perception.
I thoroughly enjoyed reading this recent paper by @yasamanbb et al (https://t.co/nU3X6KW3pT) that derives analytically why certain latent variables must lead to geometry in word embeddings. (getting Fourier modes even with open boundary but exponential kernel is neat!) I think it would be great to compare this to some of @prfsanjeevarora et al's work on this (eg https://t.co/UpK9DTEC03)
More broadly, I have been thinking about the right data generating process for language. For vision, we have latent spaces with great manifold structure (eg the SO3 pose of an object) and nonlinear mixing functions. But for language? Are there really any continuous latent variables? What is the "DSprites" of language? Is it all just co-occurrence stats or is there something more in LLM word embeddings?
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
A new paradigm & member toward 1-step & e2e generative modeling!
Great work by @Goodeat258 Mingyang!!!
cannot be more excited to read
me: learning to drift with my spindrift.
https://t.co/uDLCWlrifM
How do we build sparsity into JEPA representations by design, while preserving task-relevant information?
Introducing Rectified LpJEPA, a JEPA architecture that learns sparse, non-negative, informative representations through principled distributional regularization. 📐
📄 Paper: https://t.co/CZtKHVvKT4
💻 Code: https://t.co/xIJMYkK9fT
📝 Blog: https://t.co/OFCl3FAmik
(1/n)
Last October, we introduced Representation Autoencoders (RAE), showing that training diffusion on frozen semantic representations works and outperforms VAEs on ImageNet.
We received many questions: Can this scale to complex settings like T2I? Do the advantages hold?
The answer is YES. 🧵
🇬🇧 We are looking for people to join the team, if facing hard problems sounds like fun, reach out! Imho it's a great opportunity to work on computer vision in real-world, high-stakes systems.
🇪🇸 Dale vente hazme caso.
El grupo de ML/AI en Indra Defense está creciendo mucho últimamente, abriéndose justo ahora nuevas posiciones en el área de Computer Vision. Trabajarías desde el desarrollo de modelos propios hasta la integración en hardware real con requisitos operacionales.
Hey everyone,
We are growing at Indra Defense AI and we are looking for people with ML and CV backgrounds, from PhD level to early-career engineers.
Desde Indra Defensa AI estamos buscando perfiles en ML/CV con diferentes nivel de experiencia, desde PhD hasta early-career.
We are a mix of people with different backgrounds, from ex-academics (assist. professors, postdocs, Ph.D.s) and Ph.D. candidates to engineers. I personally believe we have a good environment, at the human and professional level.
Can ultrasound make you smell things that aren’t there?
Turns out, yes! We reliably triggered distinct scents like a campfire burn or a garbage truck by targeting our brains with ultrasound. To our knowledge, this has never been done before, even in animals.
This may be a promising modality for writing to the brain non-invasively:
🧵 1/
🚀 Our new preprint is out! We show that protein language models can predict protein-protein interactions by jointly encoding protein pairs, leading to significant improvements in PPI prediction.
https://t.co/i81iRVZAls
Really interesting report here:
1) World opinion is polarizing on China/Russia vs US lines
2) Favorability toward Russia/China is much more correlated with social liberalism than it used to be
3) Social liberalism has surged in high-income democracies
https://t.co/XwvTmKD4N5
Hemos doblado los ingresos por turismo en 10 años. Se dice pronto, pero son 60.000 millones entrando cada año de más.
El modelo productivo solo cambiará cuando esta línea deje de subir.