"Explainer videos. The output format I am most bullish on"
Claude explainer videos have been cool for exactly 1.5 days. Now I already can't stand them and block any account who slide a claudeslop video my tl.
We're starting to leave the territory where you'd test an LLM by e.g. "create an svg of pelican on a bicycle". As one idea to generalize it, I was interested what Opus 5 would do if I gave it the first paragraph of the Lord of the Rings, a 1M token budget (~$10) and asked for three js render of it. Opus went off for ~2 hours and wrote 5500 lines of code that (procedurally) rendered the story. It's kind of janky but fun. But it's a bit mindboggling that the LLM has to place and orchestrate various polygon assets in (x,y,z) coordinates and write code that animates it all, and that it even does anything at all.
I also like this kind of examples because no one in their right mind would ever spend the time to write something this custom but LLMs have all the stamina and patience in the world, so it's an example where we go from "no one would ever do this" to "sure, why not, it's ~free". There might be a lot more. But I'm excited about creating hyper custom worlds that you can imagine dropping players into, e.g. here to participate in the LoTR story as a spectator NPC, or one of the characters, or etc. Something like an ephemeral GTA of X on demand.
Last thought is that the domain of worlds/games exposes a weakness in LLMs: they can't easily audit their work because they aren't able to efficiently and natively perceive videos or play games within them. Here, Opus 5 had to very slowly and painstakingly take screenshots at different points, and it messed up a few times and created a bunch of jank. An example of raw capability (multimodal, gameplay) that I think is still quite lacking.
@Dan_Jeffries1 the closest thing we have to magic is the Moore's law
if we had current models just 10X better, 10000x faster and 1000X cheaper (in 26 year if Moore's law holds). Utopia is not unthinkable.
@MLStreetTalk I have a nice frontier task (create a synthetic dataset that rely heavily on visual/agentic reasoning)
astra still best
opus 5.5 very good 2x faster 4x less expensive
sol mediocre and horrible value for money
luna, sonnet etc.. useless crap that just can't do the task
@joelgrus@bechhof@MelMitchell1@BlancheMinerva it does, they are much more like alphastar/alphago than LLMs
Environment is your computer
Model of environment is the context window
Transfomer is the policy network, just initialized with some LLM weights.
produce actions that changes the environment states
Loop
@Neryssette Mais pourquoi il prend pas 2 minutes pour mettre dans le prompt de pas utiliser les maniérismes IA les plus lourds "Pas parce que... Parce que...?" à moins qu'il trouve ça premdeg hyper stylé ?
@Geometriquement Les "telles conditions" en question : utiliser effectivement Pangram et pas les faux sites d'arnaque attrape-boomers des "liens sponsorisés" Google.
Mathematician Alain Connes, on the motivation for creating the Millenium Prize challenges:
'In contrast to Hilbert’s speech — and reflecting concern in some quarters that the CMI might be attempting a similar task — those responsible for the prize emphasize that they have no desire to determine the direction in which mathematics moves forward.
Rather, they are trying to recreate excitement about the activity of mathematics, both in the general public and among school students. “Sometimes people have the wrong idea of maths and think that it will be overtaken by computers,” says Connes. “The seven problems, each selected by top specialists in their respective fields, are totally inaccessible to computers.”'
Computers have come a long way.
Source: David Dickson. "Mathematicians chase the seven million-dollar proofs". Nature; London Vol. 405, Iss. 6785, (May 25, 2000)
deepseek : too progressive on novel algorithms, any random tricks will make it to the next run, not enough care on data
anthropic : too conservative on novel algorithms, prefers to bet on data quality
openai : just the right amount of care for data quality and novel algorithms
Your rate limit sucks because Anthropic and OpenAI are mobilizing tens of thousands of GPUs in an attempt to prove or disprove Riemann hypothesis before the other does.