i made an indie game: Introducing Moonlit Grid
A calm city-building game with endless, unhurried gameplay and a world tour
Free demo: https://t.co/aRB6ulZaCe
Full game on itch io: https://t.co/miiWa5EgzN
Today we are releasing DolphinBench: Mapping the Pareto frontier of agent memory.
Coding, tool use, and long-context retrieval already have strong public benchmarks. Memory still mostly gets graded as a quiz: ask a question, judge the answer, maybe report precision and recall.
A lot of those boards are getting saturated. They also skip the two questions that matter when you pick a memory system to ship:
At what cost?
At what latency?
Agents are not taking memory quizzes but they take actions. The fact that matters is usually missing from the request. Either it shows up in the tool call, or the agent does the wrong thing while sounding fluent.
DolphinBench grades that action after long history, and puts accuracy, cost, and latency on one board.
https://t.co/jr7WAaDPe5