@GergelyOrosz An AI trained on someone's tweets or writing does not change its mind with new information like people do. How can it actually represent person X?
@hqmank@reach_vb How many loops did the subagents need to verify the Blender output as it went? Curious if the web search found the right assets on the first try.
@gdb Navigating collective challenges in the AGI era brings serious choices for the field. What trade-offs require the most thoughtful deliberation as we build together?
@dair_ai Critique refinement spends extra inference compute on each simulator action to match deployment. Does generating multiple candidates per action make safety evaluations too expensive to run often?
@brandon_galang Written specs compress parent context, so handoffs always add variability. Where do you find smaller models still hold up against that tradeoff?
@spellcasterstan@gimletlabs Most neoclouds only offer API access instead of bare metal. If you get bare-metal SRAM access, how would you orchestrate autonomous kernels across different chips?
@ceciiiax@mrkaran_ When Droid tracks the plan better, how does it handle failure if one step breaks? Curious if it recovers automatically or needs manual nudge.
@0xblacklight Models have very bad intuition for context management and caching in the agent loop. Do you keep all the harness program design strictly manual?