I'll keep updating this thread to share some findings. The whole game play may take days/weeks (even it's a speedrun). Since Astra is expensive now, I'm not sure if I can afford that.
DM me if you think this is interesting, especially if you can sponsor me somehow. (8/n)
When Fable 5 was released, Pokemon FireRed vision driven play-through showcase had been published by Anthropic https://t.co/BQ0RkXqXQh . I have tried to reproduce that in GPT 5.5 and 5.6. They don't work very well. Since Astra is out, I start trying it. And...
it works! (1/n)
Another issue of GPT 5.x is unncessary situation querying. Since one game play may take many minutes. And it is a long horizon task. The model tends to query and describe the current game scene again and again. Which just takes a lot of time.
Astra does not have this issue (at least I haven't observed yet). It always provides neat response and pushes game progress gradually (5/n)
Astra is much better. It can do more accurate path plan (converted to key commands). And when the player is blocked unexpected, it can find a way around without looping. All these help the model to move from one place to another more efficient. (4/n)
One issue is the vision understand ability. The player avatar occupies two cells. Human can easily understand which cell is on the left/right/top/down side of it. But the GPT vision sometimes fails to do that. So it may take a long route or several tries to get to the target place. And when it's blocked by some obstacle, it may fail to understand what the obstacle is. And it may start wandering in some loop. (3/n)
So I just find some speedrun guide online. Then convert that to some milestone based instructions. Then setup some harness using mgba emulator (pure vision with no memory access). Then use Codex to do agent play.
GPT 5.x can execute the workflow. But there are some major issues causing it to not work well. (2/n)
The concept of closeness, which means the distance between local optima of training and task distributions. If we optimize this then it would be possible to achieve better loss on OOD tasks while having the same pretraining loss. Maybe a bit close to meta learning? As the closeness relies on the task distribution it depends on mixing of data sources during pretraining.
Qwen 3.8 Flash Next is releasing Tomorrow. 125B paramters +51B N-gram and 6B active. Its based on the next generation Qwen 4 architecture.
Qwen 4 is coming
Have you read blog posts from @Jianlin_S (top talent in Kimi who invented RoPE). There are a lot of theories behind linear attention. And KDA is not just state+forget mechanism. Strongly recommend every post there. For example:
https://t.co/NsVmqtAEeP
https://t.co/IHMhMlCkOw
(those are in Chinese though)
The mighty A100 fleet are mission-capable from 2020 through 2029. NVIDIA computing is more than chips. CUDA gives developers and NVIDIA engineers a common platform to continually upgrade Ampere, Hopper and Blackwell throughout their useful lives.
CUDA makes NVIDIA computing versatile. Versatility makes it fungible. Fungibility drives utilization and extends durability, making NVIDIA compute a productive asset: rentable, durable and financeable.