Deciphering the "black-box" nature of LLMs.
Hi everyone, today I'm sharing a machine learning research paper I’ve been working on in the field of Explainable AI, specifically Attributive Reasoning for interpreting Large Language Models using Reverse Markov Chains. Full article: https://t.co/peItAo8NcN
[Open question] Why do LLMs "cheat" to beat benchmarks instead of just doing the task? My guess: we reward them almost entirely for the final answer, not for how they got there. So a shortcut that scores well is, to the model, just a smart move. Right result, wrong "means" (a.k.a. reward hacking). We design loss functions and RL to reward results over almost everything else. What if evals also looked at the means --the actual steps and tool calls an agent took -- not just the output? The way I'm thinking about it: instead of trying to list every possible cheat (which is near-impossible), you define a few simple rules the process must never break, no matter what path the agent takes. For example: - it didn't touch anything outside its sandbox - it didn't edit the tests or peek at the answer key - it only used information it was actually given everything it claimed to do, it really did (no fake "done!") - its answer traces back only to sources it was allowed to use A cheat has to break at least one -- even a cheat we never thought of. I come from interpretability and agent-safety research (independent), not RL or post-training -- so I'd genuinely love to hear from people who live in that world. What am I getting wrong?
@sama@elonmusk@janleike@sleepinyourhat@mia_glaese@AnthropicAI
"They called us in and wanted us to stop what we were doing."
Three friends noticed higher learning hadn't changed in centuries, so they made their own fix.
At first, their university told them to stop. Now, they’re partners.
This is Klar. Built by Isabel, Andreas, and Eric.
Excited to share that RecruitBase is in Microsoft for Startups.
Every company is about to deploy AI agents. Almost none have a real way to evaluate which one to actually trust in production.
We're building the trust layer for Agents in production. Starting with AgentFit.
https://t.co/T0OCfnx37i
@contextconor@hyperspell True leadership🙌🏽. Caring is also a productivity multiplier, people do tend to do more and feel more motivated when they feel cared for. Great work @contextconor and the @hyperspell team
I hate copy-pasting from ChatGPT every five minutes. You start, you can't stop, and eventually your own voice goes missing.
So I'm building a creative writing tool that actually works:
— moods that match how writers actually think (Blank Page, Messy Middle, Flow, Editing Grind)
— AI only when you summon it, never auto-inserted
— fact-check any sentence with one click
— dictate when your hands can't keep up with your thoughts — write in your characters' voices
Made for:
— college students
— screenwriters
— poets
— bloggers
— academic writers
— anyone who's wanted to write but hasn't started
https://t.co/eqrZUlpYrC
Genuinely want feedback, even the harsh kind.
I hate copy-pasting from ChatGPT every five minutes. You start, you can't stop, and eventually your own voice goes missing.
So I'm building a creative writing tool that actually works:
— moods that match how writers actually think (Blank Page, Messy Middle, Flow, Editing Grind)
— AI only when you summon it, never auto-inserted
— fact-check any sentence with one click
— dictate when your hands can't keep up with your thoughts — write in your characters' voices
Made for:
— college students
— screenwriters
— poets
— bloggers
— academic writers
— anyone who's wanted to write but hasn't started
https://t.co/XUDZo9tfgH
Genuinely want feedback, even the harsh kind.
Built a contract analyzer for the non-legalese: https://t.co/fDsioH8a2O
Upload your contract and get back:
built a contract analyzer in a weekend to solve my own personal problem
Hey everyone,
I kept signing freelance contracts without fully understanding them. Not because I didn't care — just because 14 pages of legalese at 11pm is genuinely impossible to parse.
So I spent a weekend building **Contractio** — you upload a contract (PDF or DOCX) and get back:
* Responsibilities of both parties
* Liabilities and exposure
* What happens if someone breaches
* Red flags and weak clauses (rated low / medium / high)
* A plain-English risk score
Every vendor claims their agent can automate everything. Every demo looks polished. And there’s a lot of pressure to “do something with AI” before competitors do.
But in practice, most businesses are stuck asking questions like:
Will this actually fit our workflow, or are we going to redesign everything around the tool?
Does it work with our existing stack?
What happens when compliance/security/legal gets involved?
Are we evaluating real capability, or just buying into marketing?
That decision process feels way too fuzzy for something that can become expensive fast.
So we started building AgentFit — an open-source way to evaluate AI agents against actual business constraints instead of hype.
The idea is simple: compare agents based on fit (workflow, infra, compliance, etc.), not just flashy features.
Still early, but I’d genuinely love feedback from people who’ve tried adopting agents in real companies:
GitHub: https://t.co/oVRshQts62
Site: https://t.co/lqiRMlXh1d
Choosing an AI agent for a business feels weirdly broken right now
Every vendor claims their agent can automate everything. Every demo looks polished. And there’s a lot of pressure to “do something with AI” before competitors do.
But in practice, most businesses are stuck asking questions like:
Will this actually fit our workflow, or are we going to redesign everything around the tool?
Does it work with our existing stack?
What happens when compliance/security/legal gets involved?
Are we evaluating real capability, or just buying into marketing?
That decision process feels way too fuzzy for something that can become expensive fast.
So we started building AgentFit — an open-source way to evaluate AI agents against actual business constraints instead of hype.
The idea is simple: compare agents based on fit (workflow, infra, compliance, etc.), not just flashy features.
Still early, but I’d genuinely love feedback from people who’ve tried adopting agents in real companies:
GitHub: https://t.co/oVRshQts62
Site: https://t.co/lqiRMlXh1d
Deciphering the "black-box" nature of LLMs.
Hi everyone, today I'm sharing a machine learning research paper I’ve been working on in the field of Explainable AI, specifically Attributive Reasoning for interpreting Large Language Models using Reverse Markov Chains. Full article: https://t.co/peItAo8NcN