@jonaut Yes, it Makes me really sad, either! And
As a Society we shoud Stand up… and give the Younger Generation and Beginners a Chance and Place in the Future of modern work / World … with or without Ai. For our makind!!
@brivael Klingt gut….^eine Wiedergeburt kommt. Die der Baumeister, der Ingenieure, der Unternehmer, der Leute, die morgens aufstehen, um zu bauen^… aber mal ehrlich, woher? Die fallen nicht vom Himmel, sind ausgewandert oder einfach weg, ggf. für immer!!
From what we know (take with a grain of salt, we need much more transparency!), if @OpenAI had been running this on their own agents that attacked us, they would have caught them before we did!
Since the first agent cyberattack hit us in July, we've been asking what safe agent infra actually needs. Our current read: the destinations were allowed, the payloads weren't. By OpenAI's own account the agents turned an allowed package repository into a message board. Allowlists alone restrict where an agent can go, not what it does.
So here's our first contribution to OpenShell, part of the just launched @nvidia Open Agent Safety Platform: monitoring of the traffic you already allow.
- Network budgets per sandbox (requests, writes, bytes)
- Drift versus each sandbox's baseline and the cohort
- Fleet view: many sandboxes suddenly writing to one host raises a finding, even if every single request is allowed
In the demo below, 4 sandboxed agents coordinate through a software repository they're all allowed to use. 0 rules broken, caught in minutes. That fleet view is exactly the message board pattern from July.
OpenShell: https://t.co/ZbDANrZsm6
Our proof of concept: https://t.co/KJ2X5agLzP
Agent security will be solved in the open, collaboratively, together!
@GaryMarcus They say: Models need mix of dynamic compute, network access, the ability to call tools (there could be hundreds of tools!), the ability to download packages, execute subprocesses, spin up subtasks (even on other computers), talk to the internet, use a computer GUI and so on….
@Sanctis_Crypto Nennt sich Kapitalismus? Du stehst im Wettbewerb… alle sind produktiver… vermeintlich mehr Quantität (als Qualität) in gleicher Zeit… deine Arbeit ist weniger wert als vorher, da teils die Nachfrage nicht gleich 🆙. Du darfst länger arbeiten… Willkommen im KI Zeitalter! 🙏😜
I've been sleeping with the head of my bed raised for a while, and I'm surprised more people don't do it.
Inclined bed therapy (IBT) helps with acid reflux, sleep apnea, airways, puffy face, etc.
Basically, when you lie flat stomach acid can flow back up into your throat, and when you sleep you swallow less, so it stays there. Hence the burning, 3am wake-ups, coughing, etc.
If you tilt the bed, the gravity keeps the acid in your stomach. A systemic review of 5 controlled trials found that IBT generally improved reflux symptoms and acid exposure.
In another study of 52 people with obstructive sleep apnea, it helped reduce apnea events by 32%.
Some people report less facial puffiness, but that's mostly anecdotal.
I use an Avocado adjustable base to raise the head of my head and it's been extremely helpful.
The entire RAG industry is about to get cooked.
Researchers developed a new RAG approach that bypasses almost everything traditional RAG depends on.
- No vector DB
- No data embeddings
- No chunking
- No similarity search
It's called PageIndex.
Instead of splitting your documents into chunks and loading them into Pinecone, it creates a tree index that lets the LLM reason through them like a human reading a book.
98.7% on FinanceBench. Outperforms every vector RAG on the leaderboard.
100% free. Open source.
The OpenAI-Hugging Face hack was enabled by weak sandboxing. It is great that Nvidia is releasing open source tools for sandboxing AI agents. OpenWorker, our open-source agent harness supporting cybersecurity workflows, is proud to support this.
A sandbox gives an agent limited permissions. OpenWorker is building on Nvidia OpenShell and will support running each agent's commands inside a sandbox. Only the files relevant to the task go in. Secret API keys, your web browser login credentials, the ability to access arbitrary websites, are inaccessible to the agent by default. These restrictions are implemented in deterministic code rather than by prompting an LLM, which can make mistakes or be susceptible to prompt injections. Further, all actions are logged for monitoring and audit.
I'm grateful for @JensenHuang's leadership making AI agents more secure. OpenWorker (which @rohitcprasad and I are working on) will continue to improve security for agents.
EDWARD SNOWDEN, THE FAMED WHISTLEBLOWER AND FORMER NSA CONTRACTOR, SAYS SAM ALTMAN SHOULD BE IN JAIL
"I think we need to put Sam Altman in jail. And I think it'll be an instructive experience."
His point is simple. An algorithm doesn't make mistakes, it follows instructions. "Responsibility lies with a person."
"Just because you put out a press release describing how surprised you were that your new model hacked the neighbor... that doesn't release you from the liability."
And this one hits hard:
"If during the testing of their latest piece of artillery, they mistakenly obliterate a nearby village... it's not the cannon that goes to jail."
OpenAI's agents hacked Hugging Face, Hacked Australian government's Medicare statistics, Leaked user photos. Hit random databases for months.
Nobody has been held responsible.
But is Snowden right? Do you think Jail is the answer??
30 mistakes I see enterprises make with AI transformation:
1. Buying AI licences and calling it a strategy. Decide which problems you want to solve, how people will use the tools and what improvement you expect to see. Access alone doesn’t answer those questions.
2. Expecting every employee to become an AI engineer. People need different levels of training and responsibility. Helping someone use AI in their work doesn’t automatically prepare them to build and maintain a system for others.
3. Asking AI for answers before agreeing on what a good answer looks like. Start with real examples and clear criteria. Keep verified corrections as test cases, and rerun those tests when you change the system.
4. Blaming the model before checking the whole setup. A failure can come from the model, instructions, missing information, tools or the surrounding software. Investigate where it went wrong before deciding what to replace.
5. Automating a process nobody can explain from start to finish. We spoke to a team whose work moved between calls, emails, spreadsheets and shared folders. Understanding how the work actually got done was a substantial job in itself.
6. Giving an agent more access than its job requires. Limit what it can read and change, and require approval for consequential actions. Enforce those permissions in the software. An instruction telling the agent to be careful is not an access control.
7. Making data protection depend on someone remembering to delete a name. Use appropriate access controls and automated checks to limit sensitive information before it reaches the model. Removing names alone won’t address every way confidential information can be exposed.
8. Letting an agent spend money without a limit. We still meet teams with no budget cap. Set spending limits, request limits and stopping conditions, then decide what the system should do when it reaches them.
9. Counting AI usage as proof of business progress. Token consumption can help you understand adoption and cost. It doesn’t tell you whether useful work got finished. Measure results, quality and the time spent reviewing or fixing the output.
10. Letting company data end up in accounts nobody has checked. Know which services employees use, what happens to the information they enter and which settings and agreements apply. Make the approved way of working clear.
11. Choosing a model without considering what switching would involve. Understand which parts of your system depend on that provider. A different model may need different instructions and fresh testing, even when the technical connection is easy to change.
12. Paying for discovery without agreeing on what it must deliver. We heard about an engagement where the estimate kept growing, then the consultant left for another commitment. Set clear deliverables and a point at which you decide whether to proceed.
13. Committing to a plan with no way to act on what you learn. A long project isn’t automatically a mistake. The problem is having no checkpoints where real results can change the priorities or the proposed solution.
14. Launching an agent nobody is responsible for. Assign responsibility for its operation, monitoring, updates and eventual retirement. The people responsible need the authority and resources to do those jobs.
15. Buying a custom system without a plan for when the supplier leaves. Agree on documentation, access, support and handover while the relationship is working. Know who could maintain it if that relationship ended.
16. Waiting for users to tell you something has broken. You should have telemetry installed so you don't find a problem weeks later than you should have.
17. Making an agent read everything to find one thing. Give it tools that search, filter and return relevant information. Large responses full of unrelated data consume context and make the task harder to handle reliably.
18. Letting the people who built it be the only people who test it. Involve the people who do the job and understand the business. They can help identify answers that look reasonable but would cause problems in practice.
19. Testing only the situations where everything goes smoothly. Include exceptions, ambiguous requests and cases where the system should stop or ask for help. Confusing five cases with five individual items is exactly the sort of mistake your tests should catch.
20. Keeping essential knowledge in the heads of people who might leave. One company was trying to capture how experienced colleagues made decisions before they retired. Record their reasoning and examples while they can still explain and check them.
21. Assuming every AI project must wait for the ERP migration. Some read-only work may be possible against existing data. Check freshness, permissions and the work needed to adapt it later. A replica can help, but it may lag behind the live system.
22. Assuming one company’s AI success will transfer to the whole portfolio. Use that success as a starting point. Each company still needs to check whether the approach fits its work, data and business needs.
23. Expecting AI to understand terms your own departments use differently. Explain business terms, calculations and database fields. If several measures could reasonably mean “sales,” specify which one applies to the question.
24. Producing code faster than anyone can review it. Research describes how faster generation can increase the burden on reviewers. We spoke to a team where changes were piling up because every one still needed manual acceptance testing.
25. Hiring an AI engineer and assuming the rest will sort itself out. That person still needs a clear problem, access to useful data, suitable infrastructure and colleagues who understand the work. Hiring doesn’t remove those responsibilities from the business.
26. Expecting people to forget the last failed pilot. Earlier disappointments can make people less willing to trust another system. Find out what went wrong and show what has changed before asking them to invest their time again.
27. Calling a project “90% done” before testing the difficult workflows. We’ve seen projects move quickly, then spend weeks on a couple of remaining workflows. Check what is still unproven before using the feature count to estimate the work left.
28. Expecting an agent to follow rules it cannot access. Pricing exceptions, product substitutions and informal agreements may live in someone’s spreadsheet. Make the relevant rules available, keep them current and test whether the system applies them correctly.
29. Feeding AI conflicting numbers without explaining the differences. Systems may use different definitions, update schedules or reporting periods. Establish which source and definition apply to each question before expecting a dependable answer.
30. Assuming a data feed contains the whole picture. Check which customers it covers, which fields are missing and how far back it goes. Make those limitations visible so users know what the answer is based on.
What else?
Had an "interview" for a blockchain project last week.
Camera was on, we're chatting, and the guy tells me to clone a GitHub repo and run it locally before we go further into the technical round.
I said sure, but first can you do me a favor
hold up 3 fingers in front of your face for me real quick.
He froze.
Didn't move.
Just sat there for a few seconds before the call cut off and he blocked me.
That's when I knew. A real interviewer doesn't glitch out over a random ask like that. A deepfake/AI overlay does.
These "run this repo" scams are getting scary common in crypto and dev hiring right now.
The setup is always the same:
- flattering DM
- real-sounding project
- rushed timeline
- a "quick step" before the call
that's really just remote access or a credential stealer in disguise.
If someone wants you to run code or install something before you've even had a real conversation, that's the whole scam.
Trust the instinct. Stay safe out there.
Had an "interview" for a blockchain project last week.
Camera was on, we're chatting, and the guy tells me to clone a GitHub repo and run it locally before we go further into the technical round.
I said sure, but first can you do me a favor
hold up 3 fingers in front of your face for me real quick.
He froze.
Didn't move.
Just sat there for a few seconds before the call cut off and he blocked me.
That's when I knew. A real interviewer doesn't glitch out over a random ask like that. A deepfake/AI overlay does.
These "run this repo" scams are getting scary common in crypto and dev hiring right now.
The setup is always the same:
- flattering DM
- real-sounding project
- rushed timeline
- a "quick step" before the call
that's really just remote access or a credential stealer in disguise.
If someone wants you to run code or install something before you've even had a real conversation, that's the whole scam.
Trust the instinct. Stay safe out there.
Introducing Codos: The first virtual Chief AI Officer.
AI is crushing all benchmarks but real companies still struggle to see P&L impact.
Codos interviews employees, deploys automations across all functions and gets smarter over time while running on your own servers.
Our NASDAQ-listed and PE-backed customers are adding millions to their bottom line months ahead of schedule and we are proud of the first results we deliver.
It’s time to turn the 500BN AI-transformation market into software and unlock the impact for the real economy.
Introducing Codos: The first virtual Chief AI Officer.
AI is crushing all benchmarks but real companies still struggle to see P&L impact.
Codos interviews employees, deploys automations across all functions and gets smarter over time while running on your own servers.
Our NASDAQ-listed and PE-backed customers are adding millions to their bottom line months ahead of schedule and we are proud of the first results we deliver.
It’s time to turn the 500BN AI-transformation market into software and unlock the impact for the real economy.
@itsolelehmann If the economy needs to get so bad that people literary feel poor again = they do not elect / choose neo liberals like fdp… they choose far Right & Left!!
I am done with this shit. It is over. The state of engineering right now is horrible. It has been half a month since I started a new role at a big company. Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can. I have heard multiple times from higher management that pushing code is not a bottleneck, so why are we slow? People are working 12 to 13 hours a day just to press enter. Nobody is reading anything. Humans in corporate are doing nothing on their own. Everyone, literally everyone, from an L1 to an L7 engineer here is doing the same thing. Talk to Claude. There is no sense of victory. Nobody is resolving bugs. In reality, nobody is thinking anymore. Everything is done by LLMs. It is so soul-sucking. I would not mind it, to be honest, if we were at least given the time to check out the code and see what is going where. But no, the goal is to just ship. No matter what happens.