very bullish on Bend, this is exactly what ive been trying to do in other languages, but you can only get so far with weak type systems and lint rules.
when you have a codebase that can formally verify itself, you can ship at an incredible pace. code review is solved
i'd argue you don't see how bad they are coding :P in fact, the failure modes are the exact same as the meeting summaries:
- miscategorization of intent, ("no i meant...")
- losing touch with the overall "plot", (horrible design)
- half-assing through long lists, (=> refactors)
- speaking their own weird dialect, (=> isRecord)
most of it comes down to herding it a little differently:
defining workflows that do not let them fall into these gaps, preferably adding something verifiable, giving more context...
realistically, they can augment most knowledge work already, the real issue is with the post above, the "we are here" is completely wrong.
all they do is accelerate your intent, sometimes in a messy and in-complete way, sometimes requiring a lot of corrections.
until you can say "claude go make money" and come back to find yourself a millionaire, you cannot realistically say they are automating engineering jobs, and this, we're not getting close, at all. (even @usr_bin_roygbiv with his 20 jobs, I'm sure, is not just pointing them at Slack and asking them to do their thing)
why else do you think anthropic/openai has mathemathicans involved with their breakthroughs if LLMs "solved" math?
@kuberdenis@AnthropicAI They are so strict on automation and external programs injecting inputs.. I’d be super worried about a ban. Really cool concept though.
These days in code review I’m not so worried about code correctness necessarily. We have good checks for these things. But what about: the change is fundamentally at odds with the design philosophy of the system? In other words: the change is “correct” in a vacuum but the ship was being steered in the wrong direction. How do you think about solving that?
best coding agent use case for Jev:
enforcing rules a linter can't
coding agents tend to break rules and some rules can't be codified
Jev can now score every turn against such rules and tell the agent what to fix asap
open-source: https://t.co/yfw9KCm2Cu
The "whiteboard defense:" I should be able to pull you aside at any moment and ask you to explain any customer-facing system you've shipped. You should be able to clearly explain how it works and defend the decisions you made. This is my benchmark for responsible AI usage.
I don't expect line-level familiarity with the code. I don't care if you remember the exact function name or implementation detail. You may not even know it. I don't care.
But if I ask "why did you do X instead of Y?", "what happens if this actor behaves maliciously?", "what data structure did you use here and why?", or "where does this fail?" you should be able to answer confidently.
For PoCs, demos, experiments, whatever: I don't care. Generate 100% of it and understand none of it. Speed over quality every time in those specific scenarios.
But if you're shipping customer-facing work, you can't be shipping things you don't understand at a high level.
Right now it’s just a small thing that a few of my immediate teammates and I are trying out. The use case is basically like this:
- we’re all leading different swim lanes/features/products which have various technical and strategic relationships
- we each primarily drive agents from our own “brain” workspaces which have all the context that we know and care about as human DRIs. This means all the repos we work in are cloned, but also meeting transcripts, strategy documents, myriad artifacts from agent journaling sessions etc are all available
- in many cases I want to ask someone on my team for something.. but mostly that means asking their agent brain indirectly (meat proxy). So, why not just cut that meat proxy step out? Let the agent context lakes talk to each other and then escalate to us when necessary based on the criteria and rules of engagement we define
We have agents on crons just waking up in intervals, catching up on latest message board state, and then pushing things forward. There’s more nuance than that, but at a high level that’s what’s happening. It’s mostly helping assuage coordination tax which is a growing problem as overall work velocity accelerates. Otherwise lots of “wait, so and so is doing X which is incompatible with the Y ADR we were banking on”
@backnotprop@dillon_mulroy It definitely feels like the future when you wake up to a slack message where your agent is asking you for your input on things that it discussed (and meaningfully moved forward!) with other teammate’s agents overnight
@backnotprop@dillon_mulroy Our team is doing similar but mostly for my agent to talk to other teammates agents to coordinate work. It’s quite nice. Basically just an irc server for now.
🚨 TODD BEAMER, UNITED 93 PHONE CALL: "LET'S ROLL" 🇺🇸
Todd: Hello… Operator… listen to me. I can’t speak very loud. This is an emergency. I’m a passenger on a United flight to San Francisco. Our plane has been hijacked.
Lisa: I understand. Can they see you?
Todd: No. There are three that we know of. They have knives — razor knives, like box cutters. Someone announced from the cockpit there was a bomb. It sounded fake.
Lisa: Your name?
Todd: Todd Beamer. United Flight 93.
Todd: They killed one passenger in first class. They forced most of us back. Fourteen of us here. Five flight attendants. The guy with the bomb ordered us to sit on the floor.
Lisa: Are you okay?
Todd: We’re going down… wait. No. We’re leveling off. We changed directions. We’re flying east again.
Todd: A guy named Jeremy called his wife. She told him two planes hit the World Trade Center. Lisa, is that true?
Lisa: I have to tell you the truth. It’s very bad. Both towers are gone. A third plane hit the Pentagon. Our country is under attack. I’m afraid your plane may be part of their plan.
Todd: Oh God. Lisa, will you do something for me? Call my wife and my kids. Promise me you’ll call.
Lisa: I promise.
Todd: Our home number is… You have the same name as my wife. Lisa. We’ve been married ten years. She’s pregnant with our third child. Tell her I love her. I’ll always love her. We have two boys — David, he’s 3, and Andrew, he’s 1. Tell them their daddy loves them and he is so proud of them. The baby is due January 12th. I saw an ultrasound. We still don’t know if it’s a girl or a boy.
Lisa: I’ll tell them. I promise, Todd.(Lisa patches in the FBI.)
Agent: Todd, your plane is on a course for Washington. Best guess is the White House or the Capitol.
Todd: I understand. I’ll be back.
Todd: Everyone knows this isn’t a normal hijacking. We have decided we will not be pawns in their plot.
Lisa: What are you going to do?
Todd: Four of us are going to rush the one with the bomb. Then the cockpit. A stewardess is getting boiling water. We’ll take them out.
Todd: Would you pray with me? [They pray the Lord’s Prayer.] Yea, though I walk through the valley of the shadow of death, I will fear no evil, for thou art with me.
Todd: God help me. Jesus help me. Are you guys ready? Let’s Roll.
SOURCE: @SterlingMSnow
@_lopopolo I got similar advice early in my career. It was framed a little differently: "get technical enough, and then the only thing that matters is agency"
As a junior engineer 6 months out of university, I got the best advice of my career: as you become more senior, people will do less and less for you; if you want a thing to happen, you have to do it.
All staff engineer career advice (bring solutions not problems, see risk from the future, effectively manage up) is downstream of this. Downstream of needing to do things to get to outcomes.
I wrote this @ 1 am because I couldn’t sleep. But TLDR; K-shaped theory of AI applications. A handful of people using it can have 1,000x the impact than the vast majority of people.
Selling the compute resources along with the AI to a few corporations/institutions will be very profitable.
“Daddy what were you doing during the singularity, what was it like?”
“Well son I was mostly reading about it on my phone and talking about it in group chats, and otherwise living my life as I ordinarily would have while my friends and family and coworkers completely ignored it”