excited to finally announce that @AmbrookAg has raised $29M in funding to date, including a $26.1M Series A led by Thrive Capital and @zoink.
thank you to everyone who has been part of our story so far. I can't wait for the story to come.
https://t.co/JjadSKWHMK
@mitsuhiko I'm curious: is it because of latency with the multimodal inputs to the LLM or just Playwright itself being slow. If it's the former, perhaps a subagent with a smaller specialized model would help?
Interactive Reasoning Benchmarks are the next step in frontier evaluations
Hear @GregKamradt share why measuring human-like intelligence requires multi-turn environments
Including a sneak peak of ARC-AGI-3
Want to help us build interactive evaluations? We're hiring
Sonnet 4: I'm in what appears to be a house interior (likely my home in Pallet Town based on the tile pattern). I can see my character (the sprite with a hat), a TV, and what looks like my mom (the other character sprite) in the room.
Realizing that none of the LLMs really understand what's going on in this image.
Qwen 2.5: "it appears you are standing next to a Trainer in Pallet Town, and he is likely wandering around. Since he's standing in front of you, I think the best initial action would be to interact..
The sheer amount of engineering here to make this all "just work" out of the box is amazing. In other news, running your own LLM inference is surprisingly simple (at hobbyist scale.)
Wrapping up my last day at @AmbrookAg after a great four years! My #1 learning was to how to have extreme agency with a scrappy team, from resolving outages to launching our first paid ads and AI products.
No next plans as of yet — excited to take some time off to travel, hack on some AI side projects and move back to NYC. Would love to chat with folks doing interesting work here!
@lateinteraction Noticed last weekend that the CoT module doesn’t take into account if a model has reasoning capabilities. Do y’all already have something planned here? I didn’t find a GitHub issue to this effect.
Kudos to the @CloudflareDev AutoRAG team - built a fully working RAG prototype in a few hours yesterday. Also handled the little things like creating a service account key inline!
Very excited to see more TypeScript based agent libraries, but @mastra seems too bloated to put in our monolith - anyone else have a slimmer alternative?