OpenClaw deleted around 400k LOC of its own tests without much change in code coverage. Modern models just love writing tests for every tiny change, even if they aren't useful. This skill helped. https://t.co/a1jSBpfpns
We’re giving away $25,000 to build with Gemini!
We have $25K in Gemini credits to give away, and we want it to go to a startup building something incredible.
Tell us what you’re building, what you’d use Gemini for, and why $25K of model usage would make a difference.
Apply here by 9/6 at 11:59pm PST: https://t.co/OKRrcZYDYj
Did you know, if you are subscribed to @Google AI Pro / Ultra that you get between $10-$100 every month of credits from the Google Developer Program?
Rolling out today, you can now activate those credits directly in @GoogleAIStudio!
These credits can also be used in Vertex AI, @Firebase and across @googlecloud services.
The WebMCP Challenge launches today.
The first 1,000 builders get $20 in AI Gateway credits, and the top 10 projects get $3,600 in Vercel credits & $600 in AI Gateway credits.
Let's ship ↓
I tracked down 17 AI providers giving away fkn FREE tokens rn 💀
Estimated ~1T token value across:
- GLM 5.3 Flash
- Claude Opus 4.6
- Gemini 3.7 Flash
- Qwen 3.8
- and more...
Bookmark this for later. Full list below 👇
#1 underrated skill for your harness?
Ponytail
- It forces your agent to think like the laziest senior dev
- It can save you HOURS of research and work
It's like the developer who looks at 50 lines… says nothing… and replaces them with one.
Works with Claude Code, Codex, Hermes Agent, Copilot CLI, Pi, OMP, OpenCode, Cursor, Gemini, and more.
https://t.co/5CxRPYLsI3
@pidotdev@poolsideai I was looking at its thinking trace and I saw it think and output the same lines (“actually, …”)
I exited the agent immediately.
The task was to set up a custom provider inside opencode and link to oh my openagent config
Don't train the model, evolve the harness.
I read a brilliant blog post from Hugging Face where they took a frozen open model scoring 0% on a hard legal agent benchmark, left its weights alone, and let an automated loop rewrite only the code around it.
That code layer is the harness, the runtime wrapper that feeds the model context, runs its tool calls, and decides when a run ends.
By the time the loop finished, the system had essentially matched Sonnet 4.6 on the benchmark's headline metric, at roughly 7x lower cost per task. Zero weights changed.
The gain existed because of where the model was failing. The judge only grades files saved in the right place under the exact requested filename, and the model kept doing the legal analysis correctly, then saving it under the wrong name, dropping it in a scratch folder, or never writing it at all.
So the 0% was never measuring legal reasoning. It was measuring the harness.
Hand-tuning that layer is slow and model-specific, so they automated it. A Claude proposer adds exactly one mechanism per iteration, and an outer loop keeps it only if it clearly beats the current best, so accepted mechanisms compound.
What the loop discovered says a lot about where agents actually fail.
→ The biggest single gain was file handling, not intelligence. An automatic step that lands the deliverable exactly where the judge expects it beat every prompt change, with zero extra model tokens.
→ Code fixes transferred across models, prompt playbooks did not. The same harness lifted a smaller model from the same family by 14 points, but the tuned prompts hurt a different model family on tasks it could already finish.
→ The harness mattered more than anything else. Same model, same judge, same tasks, and five different harnesses scored anywhere between 3.5% and 80.1%.
The gains do eventually flatten, and the remaining misses look like real capability gaps. At some point the wrapper runs out of tricks and the model has to carry the work.
But the lesson holds. A benchmark score measures the model and its harness together, and until the harness is fixed, it's impossible to know which one failed.
I highly recommend reading this: https://t.co/3ZIeKhsngn
I also wrote a deep dive on agent harness engineering a while back, covering the orchestration loop, tools, memory, context management, and everything that turns a stateless LLM into a capable agent.
The article is quoted below.
BOOM!
Meet the open source Cambrian Explosion of repulsion of Anthropic!
Meet Qwythos 9B, a Qwen3.5 based GGUF that's both uncensored and quantized for efficiency.
I am running it now and it is brilliant!
A model that can reason through 1 million tokens of context, understand images and text, and even call functions.
Come and take it!
https://t.co/UFU3fas9OD
UI Skills es un directorio de skills para tu IA, para que deje de crear diseños genéricos y aburridos.
Seleccionados a mano y con animaciones:
→ https://t.co/eHY4XFWrcs
Creator of Claude Code:
"At Anthropic, almost 100% of our engineers are running 100+ agents with self-improving loops
self-improving loops help agents become better with each run."
in a 1-hour podcast, Boris explains how they build agents loops from sratch.
Claude + loops + routines + dynamic workflows - that’s the secret.
Watch the talk, then read how to apply the same playbook to quant trading below.
Anthropic research lead:
"99% of our engineers are running swarms of 300+ self-improving agents.
close the agent loop. Give the model a way to verify its own output"
in a 20-minute session, Anthropic team member explains how to build a model that improves itself.
Claude + loops + plan mode + dynamic workflows -that’s the secret.
Watch the talk, then save the playbook below.