@dillon_mulroy@uwukko Currently working on an ultra token-efficient browser tool. Basically it splits the page into sections and summarizes with a cheap model so that the main agent can choose specific page regions to look at.
https://t.co/xIjMiTPj1G
Computer use models shouldn't learn from screenshots.
We built a new foundation model that learns from video like humans do. FDM-1 can construct a gear in Blender, find software bugs, and even drive a real car through San Francisco using arrow keys.
@dwarkesh_sp This has been bugging me for a while too.
I've written some more about why computer use is so clunky and explored a few solutions recently: https://t.co/Gdttd3OGdi
@tli104 Nice work! Seems sensible to have this by default. I find agents perform much better in my use when I do manual compaction at good moments.
Another thing I do quite often is rewind + summarize from here. Would be interesting to also allow the LLM to compact the past N turns.
@fa1zvn Cool project. I think you might also like our recent paper where we create a harness that makes browser control much easier for smaller models.
https://t.co/Ih5O07RcOA
@jasondeanlee Agreed, feels like current models should be overkill for this but they still get tripped up sometimes. I'm more optimistic about smaller models with specialized harnesses handling these kinds of tasks https://t.co/Ih5O07RcOA