Model held fixed, only the context varied. Context quality predicted hallucination resistance, injection resistance, and tool use.
The model was never the variable.
If you are in a position to host your own models it would be prudent to do so, even if this means that only some percentage of your capacity to start. Your knowledge and decision-making infrastructure should remain under your full ownership and control.
Enterprises spent twenty years getting human access control right. Agents are getting broader access in their first month than any employee earns in a career, and it's enforced by a system prompt.
Zero trust for agents will be the whole security industry by 2028.
@NotionHQ The interesting bit here is the gap between 31% and 10%. L1➡️L2 is a tooling upgrade. L2➡️L3 is where the agent has to act on your systems without a human sanity checking every answer/action. The maturity ladder is tacit knowledge + process data being available to agents.
Absolutely brilliant from Palantir:
"Tokenmaxxing hijacks your value orientation and decreases your institutional fortitude and intelligence. The pursuit of high token usage incentivizes disposable scripts over robust software — with the addictive feeling of false progress. There is a reason why those selling tokens refuse to charge based on value."