We analyzed Kimi K3 vs. Claude Fable 5 for software engineering tasks using DeepSWE.
Kimi K3 gets you the same performance as Fable 5 at ~35% of the price, and it actually pulls ahead at higher pass@k's. More insights in the thread!
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5.
The full model weights will be released by July 27.
Congrats to the @Kimi_Moonshot team on this major milestone!
"The #1 reason AI SDRs and sales reps, trained correctly, actually work?
Because they'll follow up with each and every lead.
Not just the best ones." with @lennysan
Snowflake CEO Frank Slootman explains why your company priorities are wrong
“I found out early on that if you can whittle things down to just one thing, you become unstoppable. Unfortunately, people resist whittling things down to one thing because it’s really hard to decide what that one thing is.”
The former Snowflake and ServiceNow CEO continues:
“People have a very easy time telling you what their top 3-5 things are because hopefully the right things are in there somewhere… I can’t tell you how many board meetings I’ve been in where the CEO puts a PowerPoint up and it’s one bullet after another listing all of the things that are their priorities. You just know that they’re going to be a mile wide and an inch deep, swimming in glue, moving like molasses. The energy is leaving my body already just watching a long list of priorities… You’ve basically devalued what you should be doing because you’re time-sharing now with all of these other things.”
Mr. Slootman urges founders to think really hard about the one thing that matters most to your business and focusing entirely on that.
If you can’t decide, just pick one:
“I like to do things in sequence. Even if you’re not sure, do it anyways. Because in the process of doing, you’re going to find out whether you’re right, wrong, or somewhere in between, and you can adjust.”
When you prioritize just one thing, things move much faster:
“Things are going to go much quicker because have a narrower plan of attack. It’s energizing. The pace picks up.”
Source: @twistartups@Jason (Jan 2022)
All of this is going to rapidly change in the next couple years because of OpenAI Ads (and eventually Gemini Ads, and other model providers). Ads are coming to AI models and most people do not expect how big and how different ads in AI models are going to be.
We have a small glimpse into this world at Lapis but very exciting for the future of adtech
I had dinner with an old school ad exec the other night. She’s about 65, been SVP level at most of the big ad agencies in NY for the last 20 or so years, friend of my moms.
We were discussing why they were all failing, and let me just tell you. The bar is so low.
A modern day ad agency would run laps around these big guys. 💡
.@JoshuaKushner has this great quote:
“If you have to choose between the most experienced person, or the most educated person, or the person who wants it the most, you always pick the person who wants it the most.”
@danawhite on taking risks and going for it:
This does seem like a regular feature but my read is it’s easy to spin up a chatbot that can respond in slack, but this is more than that.
It’s hard to spin one up that has context channel by channel, permission, channel by channel, your whole code base and whole org context built in. 🚢🚢🚢
The basic idea is easy and v0 is a hackathon project. The product here is a lot closer to *it actually works*, for enterprise grade deployments, and after quite a bit of internal experimentation and iteration. It’s kind of hard to describe other than (per the post) it’s writing majority of code, it’s deeply integrated, multiplayer, and it starts to feel like everyone is a manager. So I understand it looks easy to dismiss on quick reading but it’s not some LLM Q&A with RAG over Slack, it’s not even OpenClaw adjacent, it’s a different way of working entirely, for people and teams. I work from Slack now.
We've kept hearing how GLM-5.2 beats Opus 4.8, and are skeptical of benchmarks - so we tested them on a real bug from the Cline repo. While both models fixed the issue, GLM was the winner in terms of cost and code quality:
- GLM used twice as many tokens (GLM 1.1m vs Opus 660K) but cost half as much (GLM $0.41 vs Opus $0.81)
- Opus finished quicker - 1.6 min and 12 tool calls vs GLM 4.7 min and 28 tool calls
- GLM cleaned up dead code and verified the build compiled before completing. Opus didn't - it left type errors that passed tests but broke the production build.
Both runs used the same Cline harness prompting and tools, so it seems GLM is RL trained to spend more tokens verifying its work before completing. Impressive work by the @Zai_org team!
The US government, citing national security authorities, has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including foreign national Anthropic employees.
The net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance.
Access to all other Claude models is not affected.
We apologize for this disruption to our customers. We believe this is a misunderstanding and are working to restore access as soon as possible.
Read our full statement: https://t.co/bwn0sximKZ
How long until every interview cycle consists of an agent going through every single post I’ve ever made on here?
I bet the top companies already do it…
How long until every interview cycle consists of an agent going through every single post I’ve ever made on here?
I bet the top companies already do it…