I'm going to show my work, build in public and share things that I think are interesting or insightful. Not opinions or criticism on events or other's people work. I want to show the process of what I'm working on. The things I learned by doing.
Help by following, sharing and replicate my work.
AgentCon Amsterdam: Beach Drinks by boxd Γ Moyai
Join the teams of boxd and Moyai for beach drinks at Strandzuid, Amsterdam. Thursday 17 September, 19:00 to 22:00. It is a few minutes on foot from the conference.
Register here now: https://t.co/eXDxNiEgdQ
We are trying out new models and inference providers for our LLM verifiers. Recently we started looking int GLM-5.3-flash on Nebius and comparing to Together.
Graph shows Pearson correlations (r) on terminal-bench 2.0
So if you want the top logprobs for frontier open-source models.
Use the DeepSeek platform for the DeepSeek pro and flash v4 models cause Together AI doesn't give them. But for GLM you need to use Together AI cause the https://t.co/3riO4bejnL platform doesn't have them.
Strange
One focus will be coding agents and the wider software factories concepts. I think that open-source is taking a beating and I am sticking my neck to fight against the slop. Design a system that can be used to gain the upper hand for the side of the maintainers, working title openMaintenance.
I'm going to show my work, build in public and share things that I think are interesting or insightful. Not opinions or criticism on events or other's people work. I want to show the process of what I'm working on. The things I learned by doing.
Help by following, sharing and replicate my work.
My goal is to have a better understanding of verification and how to use it to do model mixing, compression, loops and recursive improvement. Tools to build better agents that are reliable, capable and robust. Which learnings and modules I will use as the open-core of our agent reliability platform Moyai.
@tanishqk Hot take but I like it spicy πΆοΈ though: I think linter needs human and agent errors? The agents.md file could also just be custom lint rules know orgs that have massive ROI on these.
@kimmonismus RFS: Hook up all the hardware you can find that people might use to run local models. Run benchmarks for all models with all harness combinations on every hardware setups and model quantization.