@bentyuto@araseb_ facts, we had better frameworks started in the 90s and threw them away because of the amount of training they would need, now look where we are with training data.
Should have stuck with the cognitive science approach..
@DaveShapi People mistake massive recall and fluent mimicry for intelligence, but the "strawman machine" behavior is a smoking gun that proves the absolute opposite: underneath the slick vocabulary, there is zero structural comprehension
@Sentdex Go look how they do on the ELEPHANT benchmark, the models are very sycophantic but that is a metric that doesn't seem to matter in the AI sector.
People say they are intelligent, but they are often confidently wrong and straight up hallucinate without people even realizing.
@FinanceDirCFO Exactly — turn-of-the-century work in cognitive science was on the right track, but it didn’t scale as fast. But now the sector patches surface errors instead of fixing structural properties. No point auditing the books when the bookkeeper is the problem.
It's the scaling era, not the singularity.
The gaps that remain: situation/semantic structure, temporal consistency, correspondence under pressure.
Agency without that structure doesn’t compound — it regresses.
@elonmusk The late 90s and early 2000s did some great work, but stopped because of training data requirements.
So then they decided to do predictive instead, and used nearly all available data in the world to train their models anyways.
@rand_longevity We are in the "Scaling Era", and to think that scaling alone will get us there is misleading. Sycophancy, hallucinations and premature resolutions are just some of the few things that we need to solve before many things can happen, there is a very real ceiling coming.
@WCNegentropy@fidexcode I agree totally, they are optimized for coherence over actual correspondence, meaning that they have absolutely no relational understanding of the semantics.
Benchmarks are binary options, abstention on the models part is considered a failure, even when "I don't know" is safer.
Benchmarks often punish abstention or “I can’t tell from here” even when that is the honest and safer output.
Real trustworthiness sometimes looks like refusing to collapse a bind or refusing to guess under pressure — behaviors that many current tests score as failure.
This is exactly what I have set out to try and solve with Barbalo, benchmarks fail sometime even when it has more end-user safety.
Treating tests as binary options, like refusing to take a side is not a valid stance, just be confidently right or wrong.
https://t.co/T48nAxAQ6E
I am both thrilled and overwhelmed at this point - Barbalo is reaching a point where I need to host as well as get test users.
I very much want to go open source with this, I will soon be serving single user tunnels for anyone who wants to test the new frontend once it's built.
Writing up the rest — agents caving to polite pressure, memory systems injecting confabulations as fact. The accumulation-vs-discernment gap is the industry's biggest unpriced risk.
I measured one piece of this: the same 7B model that flatters wrongdoers 41% of the time (matching published frontier-LLM rates) does it 7.5% of the time inside a discernment layer I built. Replicated, blind-judged, independently verified.
The biggest problem with today’s AI agents isn’t lack of data.
It’s that they’re optimized for coherence, instead of correspondence, tactical fixes and noisy lessons, but rarely develop real structural understanding.
That’s why most agent loops eventually stall or regress.