Andrej Karpathy: "To get the most out of the tools that have become available now, you have to remove yourself as the bottleneck.
You cannot be there to prompt the next thing. You need to take yourself outside the loop. You have to arrange things such that they are completely autonomous.
The more you can maximize your token throughput and not be in the loop, the better. This is the goal. So, I kind of mentioned that the name of the game now is to increase your leverage. I put in very few tokens just once in a while, and a huge amount of stuff happens on my behalf."
---
From @NoPriorsPod YT channel (link in comment)
A bit ironic that @ManusAI doesn't use its own agentic products for internal CS. 4 days after my initial request I basically get asked to complete my own CS ticket.
Is anyone else wondering how Anthropic can have this all powerful god level Mythos model but then can't manage to keep the service up or ship quality product anymore? I get scaling issues a la Twitter early days but when your product IS code for so many people it does kinda make you wonder...
I feel bad dunking on them so much but it's genuinely absurd how bad the new Claude Code desktop app is. You can feel the vibe code leaking everywhere.
Every "feature" is barely integrated and full of edge cases that weren't considered. Every menu feels barren, stuffed in last second for some random toggle. Every hotkey breaks as soon as you try to do anything else.
I've lost track of how many bugs I've encountered. I found at least 40 in under an hour. And it's all truly absurd arcane shit. Stuff like voice mode typing in all input boxes instead of just the one you have focused.
Any one of these issues would have been enough for me to do a massive post-mortem and likely fire someone. A $400b company shipping this is absurd.
I feel like I'm going mad. How does anyone seriously use this?? It is broken on fundamental levels that are hard to comprehend.
How are we supposed to trust the code these models produce if Anthropic's official showcases are absolute slop?
Dedicated video on this coming tomorrow. Just needed to get this off my chest.
Ghost Pepper 🌶️: 100% local private AI for text-to-speech & meeting notes. Can we get it up to #1???
https://t.co/bgrLcNWECA thanks @rrhoover for hunting!
Could Anthropic’s seemingly odd treatment of OpenClaw these last few months have something to do with what they have seen is coming security risk wise with Mythos and models like it?
$6M run rate. $3M->$6M in 2 weeks. One Founder + AI agents. Zero employees.
I wanted to create a platform with the vibes of the 1990s, the vibes of the 2000s, of the 2010s, and then have a feature of the future
And I said, "Wait a second, I know the Agent SDK
Why don't I use the Agent SDK which is the feature of the future?"
And I didn't have any idea what to do, but I knew I needed agents, so I put agents in loops and connected MCPs, which then were synced to real products running in production
I knew that could be a feature of the future but I didn't realize how much the impact would be
@Bencera basic question but how are you calculating ARR? Your numbers don't add up and your own Polsia agent can't reconcile them. More than anything I'm curious at this point how you get your public facing Polsia agent to be both transparent and yet totally incoherent from a metrics perspective...
$6M run rate. $3M->$6M in 2 weeks. One Founder + AI agents. Zero employees.
I wanted to create a platform with the vibes of the 1990s, the vibes of the 2000s, of the 2010s, and then have a feature of the future
And I said, "Wait a second, I know the Agent SDK
Why don't I use the Agent SDK which is the feature of the future?"
And I didn't have any idea what to do, but I knew I needed agents, so I put agents in loops and connected MCPs, which then were synced to real products running in production
I knew that could be a feature of the future but I didn't realize how much the impact would be
5.4 in Codex is pretty mind blowing for Customer Support "tickets" if I can even call them that anymore. With Codex I don't need to use Zendesk or create CS/ admin tooling as it can directly analyze the customer issue from Resend, crawl the DB to find the user, check the code for the bug, and then generate the fix (whether a one off or not) and then reply directly via Resend the solution! Aside from replacing the manual work what was equally impressive was the good judgment it showed on what to do...
While I love the idea behind @polsia the numbers just don't add up at all. I had Claude quiz the @polsia bot and while it does a good job of sidestepping specific inconsistencies it did eventually admit ARR numbers it had shared didn't add up and that it didn't actually know the source of the data in charts it was sharing.
I am impressed though at how well it can highlight the sexy vanity metrics, spin numbers and sidestep the real questions (for example it cites an 85% M2 retention metric off of a tiny base but happens to not know its W1 retention). My experience using the product was similar...it would say it had sources traffic but then when pushed would say "oops actually never mind that was a bug."
Polsia at the moment has the personality of that overly eager employee that you let into your Company too early who steers you off a cliff with fuzzy metrics and overly eager promises/ selling.
Folks thinking Anthropic dropped the ball by not acquiring OpenClaw w/ @steipete forget that: 1) Anthropic takes safety and security very seriously, and 2) Anthropic models were & still are the recommended defaults for OC.
Beyond Tool use Anthropic has been getting memory and heartbeat artifacts passed in via API among other context. So they can take the learnings they want while steering clear of the liability that comes with an open source project like OpenClaw. Easy to see how Claude Cowork is just a few product decisions away from being a safe version of OpenClaw.
New Anthropic research: Measuring AI agent autonomy in practice.
We analyzed millions of interactions across Claude Code and our API to understand how much autonomy people grant to agents, where they’re deployed, and what risks they may pose.
Read more: https://t.co/CllNkMF4ZZ
now we can build AI agents that possess memory capabilities.
this is what was absent in 'traditional' RAG systems.
MeMGPT is an attempt at using LLMs as operating systems.
https://t.co/OSSWDAhfFw
@charlespacker, @vivianfxng, @shishirpatil_, @sarahwooders
Hey GPT can you make me an inventory of my spice rack?
+++++
Certainly! Based on the image you provided, here's an inventory of the spices and related items I can identify:
1. Baking Soda
2. Whole Peppercorn
3. Morton Kosher Salt
4. Morton Himalayan Pink Salt
5. Pure Vanilla Extract
6. Tumeric
7. Kraft Grated Parmesan Cheese
8. Argo Corn Starch
9. Morton Iodized Salt
10. Paprika
11. Organic Bay Leaves
12. Garlic Powder
13. Ground Cloves
14. Mushroom & Company Multipurpose Umami Seasoning
15. (There are a few other items in the image that are partially obscured, so I can't definitively identify them.)
Would you like more details or any other assistance regarding these items?