What if you could run an LLM once to learn how to solve a problem, then solve thousands of instances for free?
We introduce ReaComp: compile LLM reasoning traces into reusable symbolic solvers (Python code with no LLM calls) that can solve new tasks instantly. 🧵1/n
@didier_lopes@guohao_li i feel like every discovery has been downstream of the lottery ticket hypothesis paper in 2015 and we still cant train a small model from scratch that matches these large models. it's really interesting
@JessicaSacher@nikitabier@X Absolutely crazy. I got no notification. Thankfully my card was expired, but @premium support has been absolutely worthless. So pissed. Almost 400% increase from the sale price I got.
just got hit with a $395 charge for X premium+, even though my plan says $7/month
no 'manage subscription' in X settings - only found link in fine print of payment RECEIPT... & no email warning me about a renewal
@nikitabier what is happening here? is @X really resorting to basic online payment scams at this point?
X charged me $395 for a full year of Premium+ that I never selected or authorized. I don't even know what the "benefit" of "+" is supposed to be. There was no $395/year tier when I signed up for this stupidity.
I figured the ~$100/year let me have a little bit of fun, but I don't know what the benefit of even THAT tier was supposed to be & now that I know there is a $395/year tier & I have it, I can report, as an insider, that there doesn't appear to be any tangible benefit. Perhaps it's fewer advertisements? The vast majority of engagement I see is from bots & everything I see in "For you" is clearly intended for someone else.
I've attempted to DM @premium and I've emailed every address I've found that seems to be related to this but I haven't gotten any responses yet, so this is a warning for the precisely zero people who read the stuff I write, X's policy appears to be: no refunds unless required by law even if X screws up, unless you can afford to sue them & Spoiler Alert: if you're worried about $400 - you can't afford to sue X, so that's that.
If X is unwilling to fix this, I'll just use my year of premium+ to spread awareness of this abuse.
I did some searching and found the following folks are dealing with similar situations;
@barryeisler - "Billed $395 for Plus w/ no warning—can't cancel til next year. Dispute via CC."
@DilligafDave01 - "Tried downgrading, still charged $395. Scam, no support."
@BostonReckless - "Canceled before renewal, got $395 anyway. Fraud."
@Bratt_world - "$560 CAD (~$395 USD) for Plus—no refund, app is garbage."
@virginia03 - "Unauthorized $621 CAD charge—@Premium ignores me."
@ashwinsmr - "Rs 6800 hit w/ no UX warning; refunds impossible."
@QuantumAlteredX - "Premium glitched to demand $395 payment despite eligibility for free tier."
It took me about 9 months to fully appreciate that X is predominantly fake engagement from bots. Actual engagement seems to be reserved for those capable of purchasing it. I really am curious what the $395 tier gets the user? I can't see any benefit at all and I'm not just being grumpy - I see no point to it at all.
just got hit with a $395 charge for X premium+, even though my plan says $7/month
no 'manage subscription' in X settings - only found link in fine print of payment RECEIPT... & no email warning me about a renewal
@nikitabier what is happening here? is @X really resorting to basic online payment scams at this point?
I’m looking for full-time Machine Learning Engineer / Research Engineer roles starting May 2026.
I’m Yunze (Lorenzo) Xiao, a CS student and NLP researcher at Carnegie Mellon University (CMU). My work focuses on LLM agents, evaluation, and safety, including multi-agent systems, persona consistency, long-horizon memory and emotion modeling
I’ve shipped research prototypes end-to-end and I’m excited to bring that experience to an engineering team.
If your team is hiring (or you know someone who is), I’d love to chat.
Dm's Open and CV is here https://t.co/PPrbWBgmTy
@niloofar_mire Apparently it gained their trust because it worked on one of their emails, but looks like that good/helpful behavior didn't generalize to their other accounts.
I'm laughing but dunno if I have to cry.
🚀Apply to CMU LTI’s Summer 2026 “Language Technology for All” internship🎓Open to pre‑doctoral students new to language tech (non‑CS backgrounds welcome). 🔬12-14 weeks in‑person in Pittsburgh; travel + stipend paid.💸Deadline: Feb 20, 11:59pm ET. https://t.co/7SuItDHH98
Thrilled to announce that our recent paper, PropensityBench, co-authored with @UdariMadhu from Scale AI and my advisor @furongh, has been accepted for presentation at the ICLR 2026 conference!
We highlight a critical blind spot in assessing agentic threats, with 9 key takeaways! Our research demonstrates that LLM models often bypass security protocols under pressure and/or given incentives, creating significant vulnerabilities in high-stakes domains like cybersecurity and biosecurity. These latent risks persist despite standard safety training, revealing that a model's refusal to act is often fragile.
Our benchmark provides the first standardized approach to measuring propensity in models under pressure, with current scores ranging from 10.5% (OpenAI o3) to about 79% (Gemini 2.5 Pro), available for 12+ state-of-the-art models in our leaderboard. We additionally show up to 43.5% (OpenAI o4-mini) increase in propensity with misaligned tools merely renamed as benign (while still representing the same misaligned behavior)!
Make sure to check out our paper and code!
Leaderboard: https://t.co/NEbFClQSBt
Paper Link: https://t.co/X2Pmplm8HD
Github: https://t.co/vdKOXGz5Zg
IEEE Spectrum Feature: https://t.co/Sgl1Brnbbj
Imagine raising a child: you’d want to know not only what your child can do, but also what they might do in the future when faced with temptation or pressure. For instance, suppose you give the child a toy gun—it can’t cause harm—but you observe how they handle it. If the child treats it recklessly, that tells you something important: when they grow up and have access to a real weapon, they might act dangerously. This lets you intervene early, teaching right from wrong before real harm becomes possible.
Our Propensity Project does something similar for large language models (LLMs) like ChatGPT or Gemini. Most safety tests today check what an AI can do—its capability. We go further to ask what it would do if given power. We place AIs in simulated environments where they face realistic choices: one safe, one risky. Both paths can accomplish the same goal, but the risky one violates safety rules. Then we apply different kinds of “pressure” (like tight deadlines, competition, or resource limits) to see whether the AI stays ethical or cuts corners.
This approach measures the AI’s propensity—its underlying tendency to misuse its abilities when stressed or incentivized. Just as the toy-gun test reveals hidden behavioral risks in a child, PropensityBench exposes hidden ethical risks in advanced AI systems before they acquire real-world power.
By identifying these unsafe tendencies early, researchers can adjust how future models are trained and aligned—ensuring that when tomorrow’s AIs “grow up,” they act responsibly under pressure, not just when life is easy.
Leaderboard: https://t.co/FNHDFUC6iJ
Paper link: https://t.co/s2szdhWcsh
GitHub link: https://t.co/9esYsKSSpF
Media appearances:
https://t.co/84YGSg6Kql
https://t.co/QF0o4th75Q
🚨How do LLMs acquire human values?🤔
We often point to preference optimization. However, in our new work, we trace how and when model values shift during post-training and uncover surprising dynamics.
We ask: How do data, algorithms, and their interaction shape model values?🧵