Amazing work! Love the emphasis on interpretable, scoped AI for science. VLM-as-a-judge feels like a promising direction — curious how it compares to human eval in these tasks.
Great to work on this benchmark with astronomers in our NSF-Simons CosmicAI institute! What I like about it:
(1) focus on data processing & visualization, a "bite-sized" AI4Sci task (not automating all of research)
(2) eval with VLM-as-a-judge (possible with strong, modern VLMs)
Anthropic just dropped a $3T nuke on every AI company.
They've officially joined forces with Apple.
Their plan?
Steal the one thing OpenAI, Google, and Grok desperately need.
Here's how two tech titans just outplayed the entire AI industry:
There are traditionally two types of research: problem-driven research and method-driven research. As we’ve seen with large language models and now AlphaEvolve, it should be very clear now that total method-driven research is a huge opportunity.
Problem-driven research is nice because you have a consistent and specific goal. The goal is usually virtuous, so it feels good to have a mission and identity. However, it just doesn’t work due to The Bitter Lesson. Basically everything in classical NLP (machine translation, summarization, chatbots) lost to simple scaling. ChatGPT is a prime example—it used nothing from chatbot research and certainly wasn’t the intended end goal of OpenAI’s 2022 research program, but was a huge hit because someone (John Schulman et al) figured out the right way to package large language models as a product.
Method-driven research feels less stable because you’re constantly searching for problems and you have to be opportunistic. But I believe AI will allow method-driven research to dominate progress in most fields of science, one-by-one. The latest method (or “hammer”), as we’ve seen in AlphaEvolve, is ruthless search and optimization against a reward function (whether this requires RL or not is a separate discussion). Things that problem-driven researchers have been trying to solve for a long time like the kissing number problem will become nails hit by the hammer. Eventually the hammer will become bigger, stronger, and more general and will hit more and more nails.
So a very important meta-skill for the next decade will be knowing how to create the right environments to use The Hammer. Ironically, the problem-driven researchers, who by definition are experts in a specific problem, are well-positioned to create these environments. If, that is, they can put down their egos and pick up the hammer.
Meet the Grok 3 family, now on our API!
Grok 3 Mini outperforms reasoning models at 5x lower cost, redefining cost-efficient intelligence.
Grok 3, the world's strongest non-reasoning model, excels in tasks that need real world knowledge like law, finance, and healthcare.
Introducing Sora, our text-to-video model.
Sora can create videos of up to 60 seconds featuring highly detailed scenes, complex camera motion, and multiple characters with vibrant emotions.
https://t.co/YYpOAcrXQ3
Prompt: “Beautiful, snowy Tokyo city is bustling. The camera moves through the bustling city street, following several people enjoying the beautiful snowy weather and shopping at nearby stalls. Gorgeous sakura petals are flying through the wind along with snowflakes.”
New Preprint📣
"Beyond the Chat💬: Executable▶️ and Verifiable✅ Text-Editing with LLMs"
Getting help from LLMs when editing a document can involve lots of manual copy-pasting.✂️📋
What if LLMs could instead suggest one-click edits directly in a text editor for your review? 1/N
Time to build more length-calibrated/conscious models, datasets, and eval metrics? Any other factors playing a role like length here? A long way to go indeed; great work by @prasann_singhal!
📣 Today we launched an overhauled NLP course to 600 students in the online MS programs at UT Austin.
98 YouTube videos 🎥 + readings 📖 open to all!
https://t.co/y7sTe2Pb83
w/5 hours of new 🎥 on LLMs, RLHF, chain-of-thought, etc!
Meme trailer 🎬
https://t.co/Okv5LPQEyE
🧵
Thrilled to share that three of my papers on text generation and summarization have been accepted at #ACL2023NLP! 🎉 Unable to attend conference in person due to visa issues, but I'll be available on Twitter and Gathertown for discussions.
Do you want a text decoding algorithm to discover diverse options and can be customized as you wish?
Excited to share our recent ACL paper on text generation, and hope to see you at the Virtual Poster Session 3, tomorrow morning.
https://t.co/ix3hIlSjAo
It’s a simple, deterministic, efficient yet effective decoding algorithm.
You can replace beam search with it and apply any sampling/reward functions on top of it as you want.
Strong results on four text generation tasks including MT and QG.
Heading to #ACL2023NLP, excited to meet folks! We will be presenting work on discourse/QUD, summarization, and probing intergroup bias, see https://t.co/tzFcjXd7FU
Tagging my students who will be there @HongliZhan@YatingWu96