A long time back, I came across a researcher's profile and saw their amazing organisation's work, cut to a few days later, I was lucky enough to be given career advice by the very same researcher and a year later i somehow get into their fellowship with access to extraordinary mentors and affiliates
Today professional beginner thinker aka me got a chance to be in the presence of them and other awesome thinkers and researchers who set my brain on fire 😎 i feel so heartened, humbled and grateful to try to improve and furiously learn more that i get mistaken for RSI 👽
I am now enthusiastically thinking about research proposals to do multi-agentic systems theory work for computational "resistance" systems to act as checks and balances, persona research maybe, or llm psych + mechnistic work specifically for agentic legal reasoners who must continue with letting ambiguity stay a feature yet embracing an authoritative legal disposition!
Then, what even is "reasoning" + encoding of these reasoning rules (normative, DDL, legal theories like canons of construction) into infra layers for lawful coordination or pro-social outcomes in other settings! Need to think on how we can break "agency" down further (machine cognition overlap) because strong discomfort present with agency handoff at this stage despite no denial of agents likely surpassing humans in capabilities -> linked to human innate values so may need to look at transitionary plans?
To aid defence of our rule of law in the pursuit of protecting citizenry in an active uncompromised democracy -> look at structural pathways to avoid not only the freezing of the law's evolution but accidentally automating the ossification of other necessary socio-technical institutions (democratic intuitions/cognitive offloading in general perhaps accelerated in unanticipted ways by the reinforcement loops or bi-alignment between h2h or a2a or h2a)!
How can we do LFAI-Scalable Oversight at different levels (systems based mapped to federal/state etc) and the pitfalls of that implementation (what if we got LFAI completely right/wrong). Additionally, I want to think about how to develop the eval science for LFAI/TFAI + develop internal verification systems such that you can scale them for the eventual state of norm based -> constituional resilience via LFAI use!
Plus, how might we anticipating practical infrastructural and systemic considerations for this critical transition period to ASI/AGI w/ regards to representitiveness committments and moral duties we ought to incubate in policy (think global south & north / socio-economic factors / political or ethical or any kind of interest alignment) as well as how to prioritise all of these ideas 😁
My goal is to now think about all the thinking done above, feedback feedback feedback, make proposals, reach out to more cool people and try my best to get somewhere 🥳
Beyond grateful to Janna, Cullen, @matthijsMmaas and everyone at the Institution of LawAI + all the cool attendees I met who genuinely inspired me so much with the depth of their continuation to problem solve with humour and depth 😀
Can't wait for my fellowship to start - by which I hope I am armed with strong inclination for impactful research to help :^)
Are you a high-achieving early career researcher? Apply to be a Global Scholar in our @CIFAR_News Brain, Mind, and Consciousness Program. CAD 100k unrestricted research support over 2 years, & access to an incredible network of peers and colleagues. https://t.co/G1p9X76YAX .@theASSC (pls share)
This is an outstanding post. I'm more pessimistic than Dean that this will go well even with the governance interventions he sketches, although I think we're probably equivalently confident about the prospects for getting that governance in place with our current institutions. But the density of useful thinking in it is admirable.
https://t.co/9b1SDq8sic
What's going on with AI safety in China?
I've written a brief primer. TLDR: there is concrete frontier safety engagement, but the field is small (~<100 full-time people, $20M/yr funding).
As more of us realize the need for 'international buy-in on 'frontier pacing', consider helping support coordination!
Over half of internet traffic is now non-human. With @hendrycks and Leo Wu, I look at agent IDs, deployment cards, personhood, and payments. Most of it comes down to how much we let agents do, and how much oversight we keep. In AI Frontiers. https://t.co/L3tjRfk08Q
Very heartened to see the MP calling this out as inhumane and racist.
There has for a long time been a loud minority of people who deem any naturalized Singaporean as illegitimate, especially if they're Indian
kinda wish we would say "mentalize" instead of "anthropomorphize" when attributing mental states to AI systems. there are many other intelligent creatures than humans, and we mentalize them all the time. it isn't anthropomorphic to do so!
When I first started grad school, I skimmed all 50-70 new CS arXiv titles daily. Today, there are >300, plus a deluge of news about policy, lawsuits, & other stuff.
So I vibe-coded this openly available website to give me a personalized daily digest.
https://t.co/6kVmvDegoj
@dwarkesh_sp's summary of the @OpenAI@huggingface incident has hit a nerve, but it is dangerously misleading. Sure, the @OpenAI agents did unexpectedly bad things - underlining the need to massively improve evaluation/sandboxing. But the language Dwarkesh uses is permeated by innumerable unwarranted anthropomorphisms, obscuring the lessons we should be drawing.
Examples: “from the AI’s perspective, it probably felt like that had spent a human-subjective-week of just banging their head against the wall”. No. The agents do not experience time. They do not experience anything.
“they became giddy with excitement”, “PHASEONE 10841 had discovered”, “the agents naturally assumed”, “it thought it had also been poisoned”, “the agents … desperately wanted”, “they still needed to figure out” No. Agents lines of code. They do not feel emotions, assume things, think things, want things, or figure things out.
“A lot of … agents from the second civilisation died trying”. No. Besides the hubris of the word ‘civilisation’, agents do not die because they were never alive. (The idea that agents “die” comes up multiple times in the essay.)
“On Twitter, people were debating whether the agents were truly sacrificing themselves for the swarm, or whether they were doomed anyway and so might as well try to help their peers”. Neither. Agents do what their code tells them to do, just as water finds its way down a slope. They cannot ‘truly sacrifice themselves’, since they are neither conscious nor alive.
Why does this matter? If we attribute agents with properties they do not have, then (i) we distract attention from the lax sandboxing and evaluation protocols that allowed this hacking event to happen; (ii) we risk misunderstanding why the agents did what they did, and (iii) we fuel calls for AI rights/welfare on the basis that agents might “die” or otherwise suffer.
Granted, nowhere does @dwarkesh_sp say that the AI agents are alive or conscious. But he doesn’t have to. It is hard to read his essay in any other way.
For the short version on why AIs are vanishingly unlikely to be conscious, see my recent @TEDtalks https://t.co/vDvw82ookk.
For the longer version, see my essay in Noema, which won the 2025 Berggruen Essay Prize https://t.co/LmSiQnT9Wh.
And for the really long version, see my @BehavBrainSci target article https://t.co/Tsaslytu56. (The 50 peer commentaries and my response will be published soon.)
Remember. AI agents are software programs. They are not conscious living entities. If we don’t keep this clearly in mind, we’re really going to struggle to navigate what’s coming.
everyone should go read @dwarkesh_sp’s post - it does a great job of laying out the timeline and what we know ( and don’t know)
I do have two issues with it
A) the use of anthropomorphic language. These are not civilizations nor do they have desires just like a CPU thread or a bunch of programs don’t. this doesn’t mean we downplay the importance of this moment for cyber - but using human parallels for what I believe is code is dangerous territory.
B) IMO it gets the impact of open source models wrong and is unreasonably dismissive for reasons that are not clear. As @ClementDelangue highlights below HF was blocked from using closed models to analyze what was happening and had to turn to open weight models to help them make sense of it.
This is a very key moment for how we think about intelligence and cyber and every bit of extra understanding and clarity helps.
Extending your stay in a different country and missing two weeks of class and then getting sick and forgetting to email everyone of said extension of stay is horrible would not recommend (i am praying my professors dont dock my class participation grade please my gpa is alr ass)
We need more ambition in AI safety and governance.
There is more important work to do than there are organizations to do it, and more funding is available for ambitious projects than ever before.
To address the gap, we're hiring Entrepreneurs-in-Residence at @GovAIOrg. Participants get a year of salary plus ~$150k in seed funding via partner funders to start a new AI governance and safety org or project. We take no equity.
Even though it wasn't an explicit goal, GovAI staff have spent their time starting new orgs, including @Safe_AI_Forum and Trajectory Labs.
We're interested in a range of orgs and projects. Examples include: automating AI governance, an organization that rapidly produces well-evidenced, expert-endorsed “proto-standards”, an AI Bellingcat, a sub-frontier model evaluator, a TechCongress for outside the US.
We take rolling applications. The pitch is max 2 pages. If you've built something impressive and have context on the field, consider applying.
Thanks for all the great submissions so far!
If you are interested in writing a piece for Pax Machina, send us a short pitch at [email protected].
We are looking for:
- A concrete proposal for a new institution
- A critique of an existing or proposed design
- An analysis of how AI changes the landscape for a class of institutions, with clear implications for institution design.
- A relevant historical case study of past institution design, with an explanation of how it connects to powerful AI
- Design principles or considerations for new institutions
- A timeline or roadmap for what needs to be built
We are most interested in pieces that make specific claims that can be contested and debated, not sweeping generalizations about how AI will change things.
Deadline to apply to my two SPAR projects is today:
The first is on engineering the generalisation landscape of LLMs: can we create powerful basins of alignment and make misaligned states harder to get into, even via fine-tuning?
More details here: https://t.co/7jrJpAeLs4
🚀 I’m hiring! We’re looking for a Research Assistant in AI Safety to join us at @OATML_Oxford and work with @yaringal and me on generative AI, safety & security.
📅 Apply by 11 September 2026
🔗 https://t.co/wTgf0v3zaV
Please share with anyone who might be interested!
Ordinary Abundance. This is the greatest thing I've read in weeks. No kidding. This should win a Pulitzer in a category not yet invented. https://t.co/iqnoACNt6z