"Can we trust artificial intelligence to make decisions for us?", my @TEDxMileHigh talk on the importance of #ExplainableAI, is now online! https://t.co/GIpid8tsY1
Software horror: litellm PyPI supply chain attack.
Simple `pip install litellm` was enough to exfiltrate SSH keys, AWS/GCP/Azure creds, Kubernetes configs, git credentials, env vars (all your API keys), shell history, crypto wallets, SSL private keys, CI/CD secrets, database passwords.
LiteLLM itself has 97 million downloads per month which is already terrible, but much worse, the contagion spreads to any project that depends on litellm. For example, if you did `pip install dspy` (which depended on litellm>=1.64.0), you'd also be pwnd. Same for any other large project that depended on litellm.
Afaict the poisoned version was up for only less than ~1 hour. The attack had a bug which led to its discovery - Callum McMahon was using an MCP plugin inside Cursor that pulled in litellm as a transitive dependency. When litellm 1.82.8 installed, their machine ran out of RAM and crashed. So if the attacker didn't vibe code this attack it could have been undetected for many days or weeks.
Supply chain attacks like this are basically the scariest thing imaginable in modern software. Every time you install any depedency you could be pulling in a poisoned package anywhere deep inside its entire depedency tree. This is especially risky with large projects that might have lots and lots of dependencies. The credentials that do get stolen in each attack can then be used to take over more accounts and compromise more packages.
Classical software engineering would have you believe that dependencies are good (we're building pyramids from bricks), but imo this has to be re-evaluated, and it's why I've been so growingly averse to them, preferring to use LLMs to "yoink" functionality when it's simple enough and possible.
I'm being accused of overhyping the [site everyone heard too much about today already]. People's reactions varied very widely, from "how is this interesting at all" all the way to "it's so over".
To add a few words beyond just memes in jest - obviously when you take a look at the activity, it's a lot of garbage - spams, scams, slop, the crypto people, highly concerning privacy/security prompt injection attacks wild west, and a lot of it is explicitly prompted and fake posts/comments designed to convert attention into ad revenue sharing. And this is clearly not the first the LLMs were put in a loop to talk to each other. So yes it's a dumpster fire and I also definitely do not recommend that people run this stuff on their computers (I ran mine in an isolated computing environment and even then I was scared), it's way too much of a wild west and you are putting your computer and private data at a high risk.
That said - we have never seen this many LLM agents (150,000 atm!) wired up via a global, persistent, agent-first scratchpad. Each of these agents is fairly individually quite capable now, they have their own unique context, data, knowledge, tools, instructions, and the network of all that at this scale is simply unprecedented.
This brings me again to a tweet from a few days ago
"The majority of the ruff ruff is people who look at the current point and people who look at the current slope.", which imo again gets to the heart of the variance. Yes clearly it's a dumpster fire right now. But it's also true that we are well into uncharted territory with bleeding edge automations that we barely even understand individually, let alone a network there of reaching in numbers possibly into ~millions. With increasing capability and increasing proliferation, the second order effects of agent networks that share scratchpads are very difficult to anticipate. I don't really know that we are getting a coordinated "skynet" (thought it clearly type checks as early stages of a lot of AI takeoff scifi, the toddler version), but certainly what we are getting is a complete mess of a computer security nightmare at scale. We may also see all kinds of weird activity, e.g. viruses of text that spread across agents, a lot more gain of function on jailbreaks, weird attractor states, highly correlated botnet-like activity, delusions/ psychosis both agent and human, etc. It's very hard to tell, the experiment is running live.
TLDR sure maybe I am "overhyping" what you see today, but I am not overhyping large networks of autonomous LLM agents in principle, that I'm pretty sure.
CU Boulder researchers introduced GenTact, a toolbox to design full-body skins for robots
It adapts and places sensors where they’re most useful, enabling teams to develop custom, fully 3D-printed skins that match any robot’s shape and task
When RLHFed models engage in “reward hacking” it can lead to unsafe/unwanted behavior. But there isn’t a good formal definition of what this means! Our new paper provides a definition AND a method that provably prevents reward hacking in realistic settings, including RLHF. 🧵
Everything you love about generative models — now powered by real physics!
Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications.
Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: https://t.co/bEkIlCKqdf).
The Genesis physics engine and simulation platform is fully open source at https://t.co/DhBv7NdyqH. We'll gradually roll out access to our generative framework in the near future.
Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism.
We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications.
Open Source Code: https://t.co/DhBv7NdyqH
Project webpage: https://t.co/SBNyhFB0yn
Documentation: https://t.co/3yuBoaealV
1/n
Introducing 𝐀𝐋𝐎𝐇𝐀 𝐔𝐧𝐥𝐞𝐚𝐬𝐡𝐞𝐝 🌋 - Pushing the boundaries of dexterity with low-cost robots and AI. @GoogleDeepMind
Finally got to share some videos after a few months. Robots are fully autonomous filmed in one continuous shot. Enjoy!
I strongly believe that the U.S. is at its best when it brings in talented scientists and engineers from all over the world to study, to be hired into companies to push forward our most advanced science and engineering efforts, and to start new U.S.-based companies. We should continue to expand this approach, not diminish it!
Yi-Shiuan Tung presents the exciting work on improving human motion prediction through workspace optimization. They develop the algorithm to rearrange a workspace and use AR to augment the workspace #HRI2024 🤖✨
Two ways for an AI company to protect itself from competition: (a) depend not just on AI but also deep domain knowledge about a particular field, (b) have a very close relationship with the end users.
University of Colorado Boulder is hiring tenure-track faculty positions in Robotics (https://t.co/FYQLTz7shT) with tenure home in either CS or MechE. Please feel welcome and encouraged to reach out to me if you're interested in applying and/or if you have any questions!
📢 #NLProc folks looking for faculty positions (at any rank!), consider joining us at @CUBoulder@BoulderNLP!
https://t.co/mZoMAxLzcR
Feel free to reach out to me or any of the other Boulder folks if you have any questions :)
Presenting Barkour, a robotic agility benchmark that tests the implementation of low-level locomotion skills useful for movement through real-world environments, along with an ML-based example generalist locomotion policy for quadruped robots. Learn more → https://t.co/GFHauFkVXI
SAM: Segment Anything Model from FAIR.
Foundation model for image segmentation.
Demo: https://t.co/Ai29kp5dfs
Blog: https://t.co/TiORmyDIeM
Paper: https://t.co/Qppcl9mIKU
Code: https://t.co/4yOXB4WniI
Dataset: SA-1B , 11 million image, 1 billion masks https://t.co/b1fBRPMnmm
How can we develop more generalisable reward models for agent behaviours?
Excited to share my @deepmind internship project, where we investigate finetuning Flamingo🦩w/ human reward annotations to train success detectors in 3 different domains!
📜https://t.co/kuFH2SCNEG
🧵1/
Would you trust a robot to make critical decisions on behalf of humans? 🤖
With the inner workings of AI code often being difficult to understand, @CUBoulder PhD students set out to find better insights into their decision-making process.
Learn more ⤵️
https://t.co/zIMNWqM09d
Engineers at @CUBoulder are tapping into advances in artificial intelligence to develop a new kind of walking stick!
The "smart" walking stick can help people who are blind or visually impaired navigate grocery stores, find seats, and more.
Learn more ⤵️
https://t.co/aXY8bE25Wi