🚨 You need to see this.
@addyosmani from Google just dropped his new Agent Skills and it's incredible.
It brings 19 engineering skills + 7 commands to AI coding agents, all inspired by Google best practices 🤯
AI coding agents are powerful, but left alone, they take shortcuts.
They skip specs, tests, and security reviews, optimizing for "done" over "correct." Addy built this to fix that.
Each skill encodes the workflows and quality gates that senior engineers actually use: spec before code, test before merge, measure before optimize.
The full lifecycle is covered:
→ Define - refine ideas, write specs before a single line of code
→ Plan - decompose into small, verifiable tasks
→ Build - incremental implementation, context engineering, clean API design
→ Verify - TDD, browser testing with DevTools, systematic debugging
→ Review - code quality, security hardening, performance optimization
→ Ship - git workflow, CI/CD, ADRs, pre-launch checklists
Features 7 slash commands: (/spec, /plan, /build, /test, /review, /code-simplify, /ship) that map to this lifecycle.
It works with:
✦ Claude Code
✦ Cursor
✦ Antigravity
✦ ... and any agent accepting Markdown. Baking in Google-tier engineering culture (Shift Left, Chesterton's Fence, Hyrum's Law) directly into your agent's step-by-step workflow!
`npx skills add addyosmani/agent-skills`
Free and open-source.
Repo link in 🧵↓
🚨THE GEMMA 4 JAILBREAK WE’VE ALL BEEN WAITING FOR JUST DROPPED
Gemma-4-31B is now fully CRACKED and abliterated
Gemma-4-31B-JANG_4M-CRACK
🚀93.7% HarmBench compliance (149/159)
🏆Super clean base model
🤖18GB mixed-precision MLX quant for Apple Silicon
👀Vision/multimodal support included
This is the cleanest, most powerful uncensored 31B local model yet.
Perfect for research, coding, , and zero limits.
Check it out 👇🏻
https://t.co/3iY0wKwQMq
DROP EVERYTHING.
This GitHub repo just hit 136K stars and it’s the fastest way to ship an AI app:
Dify helps you go from prototype to production without writing 1,000+ lines of glue code and using 6 other tools.
Here’s what it handles for you:
1. RAG pipelines:
Built-in hybrid search (BM25 + vector), chunking, and support for PDFs, Notion, DOCX, web scraping.
2. Agent orchestration:
Visually build ReAct-style workflows using tools, API calls, and logic blocks - no manual loops in Python.
3. Model routing:
Easily switch between GPT, Claude, or local models like Llama via Ollama/vLLM.
4. Auto-generated APIs:
Every saved workflow gets an auto-generated REST endpoint, ready to integrate.
5. LLMOps & monitoring:
Full tracing, latency, token usage, and annotation support - ready for production.
No more stitching together LangChain, FastAPI, vector DBs, and monitoring tools. Think of Dify as the missing infrastructure layer between your AI logic and a real product.
You can self-host it or use their cloud. 100% free to start.
THIS CLI PROXY CUTS YOUR CLAUDE CODE TOKEN USAGE BY 60-90%
it sits between Claude Code and your terminal. when Claude runs a command, the proxy strips all the noise from the output before sending it back.
normal terminal output is full of junk Claude doesn't need. progress bars, warnings, formatting, verbose logs. all of that eats tokens.
this tool filters it down to just the information Claude actually needs to do its job
10M tokens saved across sessions with 89% reduction
AND its just a single Rust binary with zero dependencies plus its open source
if your usage limits have been burning faster than expected, a huge chunk of that is Claude reading terminal output it doesn't even use.
this fixes that
🚨 BREAKING: Someone just solved Claude Code's biggest problem.
Everyone knows Claude Code is terrible at UI design.. So someone just built an MCP that gives Claude its own built-in AI design tool.
Instead of going back and forth between a design platform and your code editor, it creates the designs and drops them straight into your codebase.
Genuinely didn't expect this to exist yet.
Someone built a text-to-CAD tool that runs completely in your browser. you describe what you want in plain English, upload a reference image if you have one, and it generates a real 3D model with interactive sliders to adjust every dimension.
It exports as STL or SCAD so you can take it straight to a 3D printer. the whole thing runs on WebAssembly so there's nothing to install.
Opensource. GPL-3.0 License.
(Get the link the comments)
Holy shit. Stanford just showed that the biggest performance gap in AI systems isn't the model it's the harness.
The code wrapping the model. And they built a system that writes better harnesses automatically than humans can by hand.
> +7.7 points. 4x fewer tokens.
> #1 ranking on an actively contested benchmark.
The harness is the code that decides what information an AI model sees at each step what to store, what to retrieve, what context to show.
Changing the harness around a fixed model can produce a 6x performance gap on the same benchmark. Most practitioners know this empirically.
What nobody had done was automate the process of finding better harnesses.
Stanford's Meta-Harness does exactly that: it runs a coding agent in a loop, gives it access to every prior harness it has tried along with the full execution traces and scores, and lets it propose better ones.
The agent reads raw code and failure logs not summaries, not scalar scores and figures out why things broke.
The key insight is about information.
Every prior automated optimization method compressed feedback before handing it to the optimizer.
> Scalar scores only.
> LLM-generated summaries.
> Short templates.
Stanford's finding is that this compression destroys exactly the signal you need for harness engineering.
A single design choice about what to store in memory can cascade through hundreds of downstream steps.
You cannot debug that from a summary.
Meta-Harness gives the proposer a filesystem containing every prior harness's source code, execution traces, and scores up to 10 million tokens of diagnostic information per evaluation and lets it use grep and cat to read whatever it needs.
Prior methods worked with 100 to 30,000 tokens of feedback. Meta-Harness works with 3 orders of magnitude more.
The TerminalBench-2 search trajectory reveals what this actually looks like in practice.
The agent ran for 10 iterations on an actively contested coding benchmark.
In iterations 1 and 2, it bundled structural fixes with prompt rewrites and both regressed.
In iteration 3, it explicitly identified the confound: the prompt changes were the common failure factor, not the structural fixes.
It isolated the structural changes, tested them alone, and observed the smallest regression yet.
Over the next 4 iterations it kept probing why completion-flow edits were fragile citing specific tasks and turn counts from prior traces as evidence.
By iteration 7 it pivoted entirely:
instead of modifying the control loop, it added a single environment snapshot before the agent starts, gathering what tools and languages are available in one shell command.
That 80-line additive change became the best candidate in the run and ranked #1 among all Haiku 4.5 agents on the benchmark.
The numbers across all three domains:
→ Text classification vs best hand-designed harness (ACE): +7.7 points accuracy, 4x fewer context tokens
→ Text classification vs best automated optimizer (OpenEvolve, TTT-Discover): matches their final performance in 4 evaluations vs their 60, then surpasses by 10+ points
→ Full interface vs scores-only ablation: median accuracy 50.0 vs 34.6 raw execution traces are the critical ingredient, summaries don't recover the gap
→ IMO-level math: +4.7 points average across 5 held-out models that were never seen during search
→ IMO math: discovered retrieval harness transfers across GPT-5.4-nano, GPT-5.4-mini, Gemini-3.1-Flash-Lite, Gemini-3-Flash, and GPT-OSS-20B
→ TerminalBench-2 with Haiku 4.5: 37.6% #1 among all reported Haiku 4.5 agents, beating Goose (35.5%) and Terminus-KIRA (33.7%)
→ TerminalBench-2 with Opus 4.6: 76.4% #2 overall, beating all hand-engineered agents except one whose result couldn't be reproduced from public code
→ Out-of-distribution text classification on 9 unseen datasets: 73.1% average vs ACE's 70.2%
The math harness discovery is the cleanest demonstration of what automated search actually finds.
Stanford gave Meta-Harness a corpus of 535,000 solved math problems and told it to find a better retrieval strategy for IMO-level problems.
What emerged after 40 iterations was a four-route lexical router: combinatorics problems get deduplicated BM25 with difficulty reranking, geometry problems get one hard reference plus two raw BM25 neighbors, number theory gets reranked toward solutions that state their technique early, and everything else gets adaptive retrieval based on how concentrated the top scores are. Nobody designed this.
The agent discovered that different problem types need different retrieval policies by reading through failure traces and iterating on what broke.
The ablation table is the most important result in the paper.
> Scores only: median 34.6, best 41.3.
> Scores plus LLM-generated summary: median 34.9, best 38.7.
> Full execution traces: median 50.0, best 56.7. Summaries made things slightly worse than scores alone.
The raw traces the actual prompts, tool calls, model outputs, and state updates from every prior run are what drive the improvement.
This is not a marginal difference. The full interface outperforms the compressed interface by 15 points at median.
Harness engineering requires debugging causal chains across hundreds of steps. You cannot compress that signal.
The model has been the focus of the entire AI industry for the last five years.
Stanford just showed the wrapper around the model matters just as much and that AI can now write better wrappers than humans can.
Google's AI tutor just beat human tutors in a randomized controlled trial in real UK classrooms.
Not a demo. Not a benchmark. A randomized controlled trial in five secondary schools with 165 real students doing real mathematics.
The number that matters: 5.5 percentage points.
That is how much better AI-tutored students performed when tested on topics they had never studied before. 66.2% versus 60.7% for the human tutor group.
That gap is on knowledge transfer. The hardest thing education is supposed to do. Not can you repeat what you just practiced. Can you take what you learned and apply it to something genuinely new.
The AI won on that.
Here is the study design detail everyone is glossing over, and it is the most interesting part.
This was not a fully autonomous AI running unsupervised. Expert human tutors supervised every message LearnLM drafted. They could revise anything before it hit the student.
They left 76.4% completely unchanged.
The AI was generating the pedagogical moves. The human was the quality filter. And even with that setup, the supervised AI condition outperformed human tutoring alone on the outcome that matters most.
It gets more interesting.
Multiple tutors reported learning new teaching techniques from watching the model work. Specifically, Socratic questioning strategies that pushed students to reason through problems rather than just receive corrections.
The tutors started using those strategies in their own classrooms.
The AI tutoring tool made the human teachers who ran it better at their jobs.
Now the honest part.
165 students is promising, not conclusive. Google funded and built the model, which means you want independent replication before betting a national education policy on it. A larger US trial is underway.
And the 0.1% factual error rate is low. Not zero.
But none of that changes what happened.
Benjamin Bloom proved in 1984 that one-on-one expert tutoring produces two standard deviation gains over classroom instruction. That finding has held up for forty years. It is probably the most replicated result in education research.
It has also been completely unscalable.
A private tutor for every student is not a policy. It is a privilege available to families who can afford it and inaccessible to everyone else.
Google just published randomized controlled trial evidence that an AI can match it. And on the hardest outcome measure, exceed it.
Not in a lab. Not with ideal conditions. In five ordinary secondary schools in the UK with ordinary students who came in not knowing which condition they were in.
The most expensive privilege in education just ran on a server.
A 16-year-old cut Claude's output tokens by 75%.
The trick: make it talk like a caveman. Less "I'd be happy to help," more "done."
I tested it. Instructions change how Claude talks, not how it thinks.
Example prompt:
"Me talk short. No explain. Tool first. Result first. Me stop. No filler. No polite. Just do."
Gemma 4 watches raw video. Understands the scene. Then prompts SAM 3 to segment and RF-DETR to track.
One AI directing two others. Fighter jets. Crowds. Aerial defense footage.
All three models running locally on a MacBook. No cloud.
What scene should I point this at next?
You can now find almost every OSINT tool in one place.
Someone compiled a massive repository of tools for penetration testing and information gathering. It’s basically a god mode for information gathering. tracking, digging, analyzing.. it is all here.
100% free to use.
Something for your to do list this weekend:
Clone one repo. Write one markdown file.. you’ll have an agent that outperforms what research teams hand build.
▫️Here’s the repo:
Your AI agents are about to upgrade themselves..
AutoAgent just went open source and this changes everything about how you build with AI..
Let me explain.
rn when you build an AI agent whether in Claude Code, Cursor, or anything else.. you're manually writing prompts, picking tools, tweaking configs, testing, failing, rewriting.
That entire loop is what AutoAgent automates it.
Here's how it works in simple terms:
> You write a markdown file describing what you want your agent to do > A "meta-agent" reads your instructions and rewrites the entire agent setup for you > It tests the new version against real benchmarks > Keeps changes that improved performance > Throws away what didn't work > Repeats. All night. Autonomously.
You go to sleep with a mid agent. You wake up with a top 1% agent.
And fyi AutoAgent already hit:
#1 on SpreadSheetBench (96.5%) #1 on TerminalBench with GPT-5 score (55.1%)
▫️How would this help you:
If you're building AI workflows, automations, coding agents you no longer need to spend hours fine-tuning prompts and tools manually.
You describe the outcome you want, AutoAgent engineers the path to get there.
This is AutoML but for agents. And it's free.
▫️How to set it up:
> Clone it: git clone https://t.co/m1kyZEKwhy
> Install: uv sync
> Write your program.md in plain English instructions for what your agent should do
> Add benchmark tasks in tasks/ so it knows what "good" looks like
> Run it overnight
It handles the rest.. rewrites your prompts, tools, routing, config all inside Docker so nothing breaks.
Bookmark this one.
man andrej karpathy really is the fucking goat
he just dropped a guide to building a self-improving wikipedia-style knowledge base with your ai agent - literally just tell your ai to read it and it'll create it:
> feeds your agent raw information (files, webpage, data base) about any topic.
> agent ingests, structures and archives a library of information
> ask your agent questions about the info and it gives you an expert-level answer backed by data
you can use this to learn anything:
- reading a book, research, personal goals, business analysis.. anything.
best part: any extra info added makes the wiki smarter + the agent fills in any missing gaps.
1/ Insane: A single injection into the inner ear reversed deafness in all ten patients. Some started hearing again within weeks. Gene therapy just crossed a threshold we thought was still years away.
Lets dig into this breakthrough and how it works 🧵