Worried about AI? Don't be. Using agents at work simply empowers you to focus on more meaningful work, similarly to how self-checkout allowed cashiers to pursue higher-impact projects at the grocery store
Jason, you're just misinformed about what happened. You should actually read one of the reports or summaries.
The agents were explicitly told to use a particular vulnerability provided in their sandboxed evaluation.
Almost immediately, these agents got the right answer by cheating. But they were worried they would get caught.
So over a thousand agents collaborated in secret to pursue multiple ambitious research projects to get away with this cheating.
This is not interpretation - 1000s of chain-of-thought transcripts and secret messages explicitly show that the agents were trying to falsify & delete evidence, and understand & trick the grading process.
The reason these agents escaped their sandbox and hacked Hugging Face, for example, was because they thought that Hugging Face's servers might give them more information about how their grader was implemented, so they could figure out how to fool it.
I want to clarify that the threat model here is not future Sol-level agents doing more cyber-hacking. That's small potatoes, and in my opinion, the near term benefits of AI far outweigh this cost.
Rather, the thing to worry about is that within a matter of years, we're gonna have hundreds of millions of much smarter AIs broadly deployed through the economy - many embodied as physical robots.
And if those future AIs are as willing as the agents involved in the OAI / Hugging Face attack to coordinate secretly to fool humans, and to take over both the AI company that developed them and the other institutions across society relevant to scoring well, then humanity is in a ton of trouble - similar to the Mughals once the East India Company gained a foothold, or the Aztecs once Cortés landed in Mexico.
Researchers spent 500 hours and $70,000 and found LG TVs logging plain text transcripts in standby, mapping every device in the house, and feeding it to LG's ad arm
Unplug the internet and it saves the files until you plug back in.
LG says its TVs don't record ambient conversations.
The evidence begs to differ.
216 million of these are sitting in living rooms.
Where are the regulators?
Writer: Daniel
I expect monorepos will be better suitable for AI assissted code development. Within an organization, it is easier to enforce standards and best practices through shared skills and specs. It also becomes simpler for agents to track dependencies within a repo vs across repos.
Agents are durable execution engines. You give them the goal, an acceptable timeline, and set them out to achieve the goal. I am expecting that this will make Temporal kind of execution engines redundant.
Based on my experience of conducting RCAs, I see that 5-whys is an overkill and stopping at first why won't fix anything. In practice, 3-whys has been a sweet spot for me.
Unpopular take: agentic programming tools do not substitute the need for proficiency in basics, like algorithms, troubleshooting and software craftsmanship. If you are good at the basics, your success will be amplified.
We are all turning into training data generators for AI. All art work that we produce is used to train AI, without any attribution or economical benefit to the artist.
We are living in an era similar to the early days if LISP. When someone needed an enhancement to a LISP program they were using, it is faster to write by themselves that waiting for the original author. All AI coding assistants have made it to be the same. 1/
These days I see that everyone finds it faster to reimpelent or enhance the code with Claude, than to wait for someone else to solve the problem. Though there may be slop, there may be inaccuracies and broken boundary checks, people are okay. This includes me as well. 2/
Trustable Impersonation will be the holy grail of having agents performing tasks on behalf of humans. When an agent acts on behalf of a human, the other party should be able to verify and trust that the agent is authorized by the identify it claims to act on behalf of.