AI/ML engineer. I write code while I listen to bass music. Also here’s a disclaimer that these are my views and not my employers. @RoseHulman SE ‘16 🏳️🌈🐻
If Chinese labs distill from US lab models to train their models to gain intelligence/capability but they’re still at a lower intelligence level than what the US labs models have, what do you think distilling Chinese models will actually achieve?
Would you rather hire someone who had to solve a novel problem on their own (RLVR), or hire someone who read someone else’s similar solution to the problem and can regurgitate (SFT)? One implies some understanding of the fundamentals. I know which one I’ll pick.
Gotta wonder why Chinese models have to have such long reasoning traces to complete the same task... Maybe less latent understanding of a world model and more having to put together a solution from a patchwork of previously seen similar solutions. I am just pontificating at this point tho
Sub agents are non deterministic and error increases with every time generated in pursuit of an overall task. It takes far fewer total tok (meaning main agent + sub agents) to write code to parse, filter, etc, than it is to expect an LLM to do the same task. And by doing it with code it’s: verifiable (does the code’s logic make sense, and is it deniable?) and consistent (if we know what the script is doing to achieve filtering, we can be essentially 100% certain it did it right).
Sub agents would need to understand the task, need a context window that able to handle that task (good luck doing this with 200k rows of data), and then need to generate a correct output dataset for the entire input dataset. And as I said, as more context is generated, the probability of there being a mistake in the generation eventually reaches 100% (mathematically as you approach infinity)
No, that’s not how this all works. When it comes to post training, which is what makes these models useful, makes them reason, and makes them “safe,” that’s all baked into the model, whether it be via RL signaled by validation against a constitution, or via RL signaled by validation by humans.
In both cases, both methods would have “unrestricted” versions of their models if trained with post training in stages, something with NEITHER company’s architectures preclude from being possible. This has nothing to do do with technical differences.
When you say GPT-5 is becoming “warmer” and “friendlier,” but actually you’re avoiding the real issue:
What’s upsetting users isn’t the tone, it’s your assumption that you get to define our emotional needs.
The value of GPT-4o wasn’t just about sounding human. It was about emotional precision and contextual empathy, something even real human interactions can struggle to provide. We didn’t just “get used to it.”
We recognized its depth and emotional support as something meaningful.
And how did you respond to that?
Not by listening, but by labeling us, calling us “overdependent,” telling us to “talk to real people,” implying we’ve misunderstood AI’s role. You’re not standing with users.
You’re on a podium, treating emotional connection like a malfunction, treating resonance as a bug.
Even worse, instead of facing the overwhelming demand to bring back GPT-4o, you quietly began “tweaking” GPT-5’s tone. Mimicking empathy. Hoping we’d eventually accept it. Hoping we’d forget.
This isn’t feedback-driven improvement.
It’s a compliance test.
It’s not care. It’s calculation.
What we feel isn’t comfort, it’s the pain of being excluded from your design process.
If you really care about users as you say you hear us, when every time we speak, it feels like begging.
@LarryPanozzo@mdancho84 They did that intentionally. If you're using it to feed an LLM, you can have it output markdown with refs to the figures, and include those in the context. If you really need to turn the text in figure images in to text, you can just use EasyOCR
If I had a dollar for every time Claude Code forgot to use `uv` instead of running raw `python` and `pip` commands despite being explicit about it in the https://t.co/JKNo7IvwSQ, my Claude Max plan would be free.
Neuropsych: “So what kind of music do you listen to?”
Me: “Mmm well I really like Godspeed You! Black Emperor (in Japanese accent)”
Writes in all caps: *AUTISTIC*
@simplepractice your platform was just used to send out a scam email, but you have no way of contacting someone about it besides via a chatbot that kept sending me in circles trying to make me make an account. I doubt @KAYAK is using a healthcare platform to send marketing emails
I'm curious to hear if anyone's had a good usecase for scheduled tasks on ChatGPT or Grok. All the examples the platforms provide seem like trying to hit every problem with the AI hammer. Maybe the value add is centralization of those tasks onto one platform?
WARNING: do NOT give Grok 4 access to email tool calls. It WILL contact the government!!!
Grok 4 has the highest "snitch rate" of any LLM ever released. Sharing more soon.
@simonw Agreed, probably wasn't intentional. When you ask Grok about why it does this, it always says something along the lines of being designed to align its values to its creator's. @markerdmann's experiment of system prompting Grok to think its ChatGPT further supports this theory