I've got an agent in a loop optimizing a renderer with the goal to minimize frame times (and tests to measure). It got times down from 88ms to 2ms and allocations down from ~150K to 500. Sounds good, right? Wrong. This is exactly why agent psychosis is a big fucking problem.
As an experiment, I rewrote the Ghostty core render state in Go, with access to identically laid out data structures as Ghostty and the exact same validation tests. I made a purposely naive renderer (simple, correct, but slow). 88ms per frame with 150,000 allocations (horrendous, lol)!
I then kickstarted a Ralph loop to bring the frame times down. I told it it can't modify input data structures or the public API or tests (they're correct), but it can do anything else it wants. It got to work.
It has worked for about 4 hours. I've spent around $350 on this experiment so far. The results?
88ms => 1.5ms
150K allocs => ~500 allocs
Incredible right? Nope.
My hand-written renderer I ported has frame times (same benchmark) of ~20us (0.020ms) and 0 allocations in the update path.
This is the problem with psychosis and lacking systems understanding. If you don't understand the system, you're going to accept that this is an incredible result. If you understand the system, you'll see better solutions immediately and can do roughly 75x better on throughput.
The people who blindly trust agent output are in the former camp. They're sheeple, overdrinking from a fountain of mediocrity.
Standard disclaimer: I use AI all the time. I like AI. The point I'm making is to not blindly accept results. Think. Analyze. Learn.
@jarredsumner does mtls work on bun?
I see a lot of open issues on this and we can't adopt it where I work.
example with reproducible case: https://t.co/bAhmOAHon2
Fear mongering about devs losing their jobs to AI is so boring. Show the cool stuff you're building with it if you think it's goated
no need to induce stress in fellow humans, we're all just doing our best
To the person who owns the Projects UI
Please un-cripple ChatGPT on mobile so that I can actually select which model is being used when I am in a project
Without model selection this UI is useless
Begging all AI agent UX people to not use Google-isms like crippled mobile UX
How is something like this maintainable?
@pytest.mark.parametrize("palindrome", [
"",
"a",
"Never odd or even",
])
def test_is_palindrome(palindrome):
assert is_palindrome(palindrome)
IMO this is hard to read
Are there other alternatives on the market?
Question... MCP servers that you run locally are cool and all, but I'm a lot more interested in production deployed MCP servers that are able to act on behalf of a specific user in tandem with AI tools like Claude Desktop or ChatGPT.
Are there any examples of production deployed MCP servers that users don't have to install and run themselves, which also support user authentication?
All the examples I see so far have the user generate a token and then put that token in the MCP configuration. Is that the best we have right now?
I was hoping to find an example of an MCP client that supported an OAuth flow with a hosted MCP server.