I write on a topic very important to me. In Part I I try to address the perception gap society has regarding women in engineering, trying to reason from first principles & brief historical context (1/4) #WomenInStem https://t.co/CGHpbfmuj9
Model upgrades ≠ product upgrades.
A “better” LLM can still break your agent: tool-use, refusals, formatting, context sensitivity all shift. If you don’t run regression evals, you can fall behind the market.
In finance terms: maintaining β≈1 takes work. Evals are the accounting system.
Read more at : https://t.co/pYHBcPdHlX
We took the simulation approach that controls fusion reactors and applied it to security.
Our Okta Log Generator creates realistic auth sequences: login→MFA→activity→logout with consistent context across events. Not random synthetic data. World models built from real security expertise.
Security devs: Test your SOC rules properly.
Enterprise: This verification approach scales to any domain.
Try free: https://t.co/y4MXZ1kgij
Rich Sutton's Bitter Lesson Isn't the Whole Story
The tech world has a favorite essay. Walk into any AI strategy meeting, and someone will inevitably reference it to justify why they need more compute, more data.
But here's how to spot a pseudo-intellectual: they never mention Sutton's equally important essay on "Verification."
https://t.co/KjzjQrQKvw
Is prompt IP?
Since the launch of ChatGPT I have been grappling with this question. I am sure there is a very nuanced legal answer to this, but I am interested in it from the point of view of how to value prompt engineering and prompt protection as a startup.
From an easy to copy angle, prompt is as easy as it gets to copy. However, from the easy to develop angle, we found that it takes a long time of methodical evaluation to get to a good set of prompts.
After a long consideration my current thinking is that prompt is akin to a recipe. Like how Pepsi and Coke have recipes that are kept very secret. Prompt should have the same treatment. The ingredients are commodity and know how to put them together quickly is also commodity. The biggest question is what ingredients to use and in what quantity?
Kashikoi is a simulation engine to benchmark GenAI Agents. They generate CPU-friendly world models that autonomously interview agents and generate deep behavioral assessments.
Congrats on the launch, @aaksham and @TimGMichaud!
https://t.co/Ma0hka6Vgc
Some “it’s-not-me-its-you” holiday cheer, merry shipmas, oops, merry christmas :P
Adapted from the book `Same As Ever` by yours truly, #hypecycle#justkidding#genai (how to deal with the crazy? coming up in the next tweet)
There are ppl. There are institutions. And then there are ppl who are institutions. Last month, several people spread across the US & India lost such a person with the passing away of Arvind Kudchadker (whom I fondly called Aajoji). https://t.co/WZ1SJ2JAeD
I think one of the UX problems with agents is that most people don't know how to articulate a task to be delegated to an agent. And many are lazy to even attempt to do it (1/4)
Whats all this fuss about Agentic AI? I recently gave a talk on what it all means. Here are the slides - https://t.co/FMNZ1Jjdgj #agentic#ai#LLMs#GPT#ml#NLP