🧵 THIS WEDNESDAY ☀️
♾️@WeOwnNet 🌐 Weekly Community Call
🗓️ Jul 22 · 12p ET
🔗 https://t.co/y9hAp2sLmU
W30 AGENDA 🚀
🧠 MetaCouncil Expansion
— https://t.co/NzRefqWWBW 🥇 P0
— https://t.co/GccFrLe6PV 🥈 P0
— Kimi K3 de-prioritized (429 errors)
🧬 Embedding Bake-Off
— Qwen3 8B vs Perplexity vs Gemini
— 30 test runs. One winner.
🌐 Infra
— https://t.co/y5OBbsN89e on Cloudflare Pro
— SearXNG migration P0
📄 AGENTS.md + SKILL.md
— Open-source agent standard
🔒 #ContextGate Protocol
— 3 strikes → defer
Live Q&A. All TEAM. All community.
Not a company. Not investors. MEMBERS own it. 🤝
#WeOwnSeason004 #CommunityCall #FedArch #MetaCouncil #FlowsBros #ResponsibleAgenticAI #BuildInPublic #WeOwnNet
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!
Engineers at Coinbase are getting closer to recursive self-improvement, or loops.
For instance, agents can now ingest and summarize customer feedback collected right in the app, prioritize bugs and features based on that feedback, draft the code, security review it, and provide it to a human engineer for final review. All happens automatically each day and it learns from the human edits to improve over time.
We're finding more areas to introduce loops across the company. Instead of prompting agents with what to do, you can increasingly give them a goal, and they bring back high quality work for you to review.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
"AI was supposed to replace human labor.
It did the opposite.
For the first time in history, humans are cheaper than software.
And AI is creating more jobs than it eliminates."
Hebbia CEO George Sivulka on what the core lessons of human management mean for agent workforces: https://t.co/lV6ctAu1Zc
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
lots of conversations about base over the last week. wanted to share my candid take after a week of listening and a lot of reflection over the last 6 months.
first off - in case it’s not obvious, the first quarter of 2026 was a punch in the face. I spent 2024 and 2025 making a two pronged bet to bring base to the world: (1) builders would unlock the next wave of crypto adoption; (2) adoption would be driven by new onchain-native social experiences - creators, content, messaging. imo we made the right bet on builders, but obviously the wrong bet on social. builders did drive the next wave of crypto adoption - prediction markets, perpetuals, stablecoins - but social was not at the center of it. in fact, the entire social side of the market that many of us had been building towards - farcaster, zora, miniapps, and yes, creator coins - disintegrated completely. I was wrong - whether it was timing wrong (is $ansem a creator coin?) or fully wrong, only time will tell, but regardless, i was definitively wrong.
the collateral damage was pretty bad! and this year has been an exercise in eating shit. we realized how our focus on social had meant that base had fallen behind in key areas that were now increasingly critical - we had perps (shoutout avantis!) and prediction markets (shoutout limitless!), but both were well behind scaled competitors. and we had a lot of room to improve in unlocking base as a platform for tokenization and payments that really worked for enterprises. people lost confidence, and CT spectators reminded me weekly of all of my mistakes as often as they could. it felt bad man, still feels bad.
but if there’s one thing i’ve learned from the last decade of building in this space, it’s that when things feel the worst, the best thing to do is just put your head down and build. so that’s what i’m doing. I refocused my time and attention back to the chain away from the app, started writing code again, shipped a bunch of stuff (azul, beryl, b20, privacy, ledgers) and questioned a bunch of my assumptions: does crypto need social to grow? does base need an app? can base be bigger than coinbase?
I thought for a long time that social was the only thing that could drive the sort of viral growth to get crypto to a billion people. unsurprisingly, I now believe that’s wrong. It’s clear that better money is more than enough - we are seeing this live with stablecoins, predictions, perpetuals, tokenization and i only expect it to accelerate. I am now focused on bringing a billion people onchain just by making global finance actually work.
on the app, my focus is on building base into the blockchain for global finance. to that end, i’ve handed the base app back to the coinbase mothership, where my now good friend @cobie will be taking it from here to make it the best damn app for onchain you’ve ever seen, including expanding beyond the base ecosystem in ways that tbh i won’t love as the leader of base.
it’s incredibly hard to grow a decentralized network inside of a big public corporation. and i feel like much of the discourse on CT over the last week is downstream of this. the following things can be true: (1) base (and i) love memes and (2) brian probably won’t ever bullpost memes on the tl (this activity is illegal once you’re over 40 years of age). it’s weird and we’re working through it as we continue to decentralize base, which has been our commitment from the beginning.
we’re going to build base into the blockchain for global finance and do everything we can to be the place that the world’s money settles over the next century. we will surely have formidable competitors (welcome robinhood and stripe!) and people may abandon our cause, but we welcome the competition and believe it’s our duty to win the respect and commitment of those who rally to our banner.
in 2026, this concretely means three things: winning trading, payments, and agents.
[continued in the reply]
I'm just going to dump my whole agentic setup out here, because I see too many people missing giant chunks of this and it's hurting them.
Here's what I have and recommend:
0. an AGENTS.md that is a router -- it sends the agent to the right skills, docs, tools
1. a standard workflow doc/skill customized to my needs ... (grab Matt Pocock skills if you don't already have something) ... I tag this in most sessions with `@/AGENT_WORKFLOW.md` and it pulls it in.
2. self-healing docs for every system, and agents are instructed to keep them updated ... I tag the ones I know I need, or let the agent find them through AGENTS.md ... I also provide a more detailed summary in the first 7 lines of every doc, so they're easily greppable to find the right thing, and this is documented in AGENTS.md
3. agents always run the app ... the agent should always actually run the app itself, and test its work and fix issues as it goes, especially if running autonomously / asynchronously
4. end-to-end tests and instructions to write more and keep up to date, and docs on how to write tests, what to avoid, and a list of all the tests and what they test in yet another markdown doc ... write and run targeted tests during implementation, improve and commit with work
5. custom linters at precommit hooks looking for any problems you run across, with `--fix` fixing the problems automatically, OR if that's not feasible, it shells out to a cheaper LLM like Composer 2.5 or Sonnet to fix the problems -- NOT just flagging them, but actually resulting in cleaned code
6. cross-agent review at each major point: research, plan, implementation, and wrap-up. I mean codex, claude, cursor, whatever -- but it shouldn't be the same model reviewing the same code. And specific docs for agent review, what to look for, how to approach it. Also, personas -- looking at the code from different perspectives, such as maintainability, code quality, security, performance, AI smells, domains (e.g. "financial services expert" or whatever) ... and each persona also "owns" a set of system docs too and keeps them up to date
7. agent traces / worksheets that track what the agent is doing each session. if the agent fails partway through, you should be able to hand this worksheet to another agent and it could finish the job. commit this worksheet with the work so it's all connected and easy to reference later (you will reference these later!!), also have the agent apply git tags that correspond to specific worksheet names so they're easy to find
8. automatic agent feedback to you at the end of the session, added to a doc that is also committed with the work, that you periodically ingest into an interactive session and improve your workflows
9. a tools or bin folder that contains python or bash scripts that the agent has skills to make to make its job easier (for example, I have an `agent_review` bash script that lets the agent kick off agent reviews via CLI without knowing each agent's particular incantations) ... docs on how to make scripts effectively, and instructions to constantly build these out more
10. periodic agent sweeps through recent commits, looking for problems / gotchas from a higher level across commits
11. a coding conventions doc that is just for specific coding conventions you want to see in the code base, your review agents use these a lot (but a lot of this should be in linters)
12. an agent loop / night shift skill for autonomous work, that lays out how the agent is to approach this, from an orchestration standpoint
13. a task queue that is accessible to the agent (mine is just a TODOS.md, but yours might be in Linear etc, with a CLI to fetch via API)
14. a periodic false-confidence test audit skill that looks for tests that aren't actually testing what you think they're testing, and that fix those
15. visual regression tests -- take screenshots, compare via tool and with agent visual review, commit with work (git lfs useful here) or at least push into the PR
16. automatic performance benchmark tests that notice when performance degrades
17. performance profiling tools that can be used by agents for targeted benchmarking, trying new techniques, comparing outputs, and comparing profiles
18. end-of-shift full validations, including running all tests, performance, agent reviews, sweeps, everything -- when you return, it's all as pristine as it can be
If you have all this, your agentic coding experience is going to be very different than dry prompting and manually guiding it toward the right thing every time.
I believe in one God,
the Father Almighty,
Maker of heaven and earth,
of all things visible and invisible.
I believe in one Lord Jesus Christ,
the Only Begotten Son of God,
born of the Father before all ages.
God from God,
Light from Light,
true God from true God,
begotten, not made,
consubstantial with the Father;
through Him all things were made.
For us men and for our salvation
He came down from heaven,
and by the Holy Spirit
was incarnate of the Virgin Mary,
and became man.
For our sake He was crucified under Pontius Pilate,
He suffered death and was buried,
and rose again on the third day
in accordance with the Scriptures.
He ascended into heaven
and is seated at the right hand of the Father.
He will come again in glory
to judge the living and the dead,
and His kingdom will have no end.
I believe in the Holy Spirit,
the Lord, the Giver of life,
who proceeds from the Father and the Son,
who with the Father and the Son is adored and glorified,
who has spoken through the prophets.
I believe in one, holy, catholic, and apostolic Church.
I confess one Baptism
for the forgiveness of sins,
and I look forward to the resurrection of the dead
and the life of the world to come.
Amen. 🙏🏼
Today's Hermes Agent Masterclass is the finale! Module 10 covers security, an important topic for running agents to ensure each agent can perform its tasks while minimizing exposure. Hermes has a ton of built-in security features. In this clip, I show how you can set approvals for each profile depending on your needs.
Grok 4.5 might be the BEST model to run inside Hermes or OpenClaw RIGHT NOW.
I've been sleeping on Grok to be honestNot anymore.
It's more than 60% cheaper than Opus 4.8 and lands around $2.49 per task versus ~$12 for Fable in Claude Code. And it's fast.
So what happens when you give Hermes + Grok 4.5 its own email, its own phone number, its own debit card, and access to every tool you use?
You pretty much get an AI co-founder.
Everything you need to know about Grok 4.5 + Hermes below.
Full episode is available to watch at @startupideaspod ( thanks @nickvasiles for coming on)
I slept on Grok. Not sleeping on it anymore.
Watch
744B parameters. On a laptop. With 25GB RAM.
Colibri runs GLM-5.2 (744B MoE) in pure C with zero dependencies. The trick: only ~40B params activate per token, so it keeps the dense part resident and streams experts from disk on demand. A single 2,400-line C file. No GPU, no BLAS, no Python at runtime.
This shouldn't work. But it does.
⭐ 2.1K #AI #OpenSource
https://t.co/W8dnCQmNnL
Follow for daily dev finds 🔔
@syedmuzayan@alighodsi Hmm … “We also find that GLM 5.2 performs extremely well.”
Considering there wasn’t another #LLLmodel mentioned seems rather clear to me.