@OpenDevLog Yes, I think it’s really interesting topic at the minute and one that many engineering teams are grappling with. I’d love to see your post on the topic and I’m currently writing one myself also.
I do think you’re spot on here .
I’m really interested in building tooling that creates leverage for individual developers and smaller teams.
Have you put into practice any thing like harness engineering or software factories to help with this?
Interested to hear your thoughts on this topic.
For me, I’ve put a lot of time and energy into creating testable reproduce and stable environments or harnesses around Codex.
Such that it can take well defined and scoped issues or tickets and work on them through completion, monitoring and observability in production with the eventually aim of tracking metrics and automatically opening new issues should bug arise or regressions occur.
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field.
I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so.
First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world.
The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability.
Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended.
I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent.
Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.)
Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration.
Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building.
[Original text (with links): https://t.co/jni2tWazAH ]
@OpenDevLog Similar situation to yourself, abandoned my account a while ago but back now. Looking to ignite/reinvigorate my personal brand/accounts and came across your post yesterday.
V. brave, keep up the good work, I’ll be following also 💪
@benjaminakar Hey Benja, one of the reasons I haven’t moved to running Codex agents in the cloud for engineering workflows is that Codex locally builds up such a good memory of workflows, how and when to use plugins/skills etc.
How/does Tembo solve for this?
That continuity is the difference for me. I stay in control of the decisions without having to manually drive every step or rebuild context. It feels less like asking AI to write code and more like running the whole engineering loop in one persistent workspace.
@thsottiaux@sama
GPT-5.6 Sol in Codex nails it because it stays with the whole job, not just the code.
I can start with an outcome, and it investigates the context, plans, makes a focused change, runs the checks, follows CI and review feedback, then leaves a clean handoff.
Meet kbd-1.0-codex-micro, built with @work_louder.
Map the buttons and joystick to your workflow, and keep your pinned chats in view.
Get yours before stock returns 410.
In engineering and technology, success isn't solely determined by technical prowess or the ability to lead a team of peers. Equally important is managing up—developing a productive and positive relationship with your managers ...
https://t.co/wv21Of8dXX
I’m hiring! If you are a software engineer and interested in joining our European R&D team here at SharpenCX please check out the page below for details. Happy to chat about this role, just reach out!
https://t.co/9pA0UUUZcJ
BREAKING: @SharpenCX Acquires @OctopusCX (aka WebText) < with UCaaS and #CCaaS companies chasing elusive “enterprise” market, refreshing to see 1 focused on helping mid-market firms accomplish a digital transformation with an agent first-focus #CX https://t.co/vAql6pGcmq
We’re thrilled to announce we've acquired Webtext! This acquisition expands our mission to empower exceptional agent & customer experiences to on-premises contact centers, allowing rapid digital transformation. Read the Release: https://t.co/DHevMzKIN9
50 years ago the @EnglandRugby that 'turned up' were halied by the Ireland team and supporters in Lansdowne Road.
Today they met at the @IrelandEmbGB to mark that famous day.
Gallery: https://t.co/8m2kTBFCwf