Context compaction has always slightly bothered me...
You give an agent a standing instruction:
NEVER deploy without running the tests.
A few compactions later and somehow we've arrived at:
"User prefers testing before deployment."
Errr... no. 😂
That wasn't a preference.
Some instructions simply shouldn't be up for interpretation.
That's basically why I built MemLock. 🔒
v0.5.1 is out.
Compaction can butcher the conversation all it likes.
Just keep your hands off the bloody rules.
I'm increasingly convinced durable instructions should just be standard plumbing in agent harnesses.
Am I missing something?
I ordered ONE 990 Pro.
They sent me TWO.
I've apparently discovered an infinite hardware glitch 😂
Naturally I'm now considering testing it with something significantly more irresponsible...
DGX Spark or Mac Studio? 👀
Choose my next terrible financial decision.
Good morning, Y'ALL!
Qwopus 3.8 27B Flash is out now! I feel like this one is the largest leap of added value we've ever committed in a fine-tune!
It's very similar to the base model at xhigh, but is significantly more efficient in its CoT, sometimes by a factor of 10 or greater. Not only is it much more concise generally, but it chooses when to use it more, where needed, and when to apply it less.
It runs exceptionally well with thinking uncapped in agentic workflows, unlike the base, especially in Claude Code!
A testament to the model's ACTUAL speed (wall time), it is incredibly usable and feels snappy running on an old V100 at only 37tps average. It is BLAZING fast on a 5090 at around 100tps with MTP!
Blessed to be here, y'all; we appreciate every one of you greatly! Big thanks to @Alibaba_Qwen and @QwenDevs for making such incredible base models for us to play with!
Please post your wins and your feedback on this one in the comments below; we're looking forward to hearing from you!
New Qwopus on a Friday? What are you building this weekend???
https://t.co/a7a1XrI7SU
People are seriously running autonomous agents with permanent root and zero gates?? 😭
This is already a massive improvement!!
but I'd harden it further: agent/job-bound grants, argument-level allowlists, binary digest pinning + fail-closed expiry.
The sudo window shouldn't become a temporary free-for-all for everything running as that user.
@OAraLabs Theyre selling it as the model is sentient, when essentially they don't you to know what's in there because it's a copy of paste of their existing model with minimal tweaks...
👀
GPT-6 Astra pricing is wild!!!!!!
GPT-5.6 Sol: $4/M input $20/M output
GPT-6 Astra: $10/M input $50/M output
That’s a straight 2.5x jump on both sides.
At $50/M output, Astra BETTER NOT JUST BE ‘better’.
• It better find a free solution for my patterned bauldness
• Fix my bad knees
• Find my missing keys
• And SOMEHOW STOP MY EX-GIRLFRIEND SENDING ME HATE MAIL every month!
Because at 2.5x the price, ‘incrementally better’ is going to BURN POCKETS very quickly!!!
GPT-6 Astra is here.
We hope it will begin to enable a new generation of entrepreneurship, scientific discovery, and building.
We believe it is the best model in the world for computer use, professional work, science, coding, cybersecurity, and more.
It took us some extra time to ensure that we could meet the safety and alignment standards required for this capability level, but we think you’ll find it worth the wait.
It scores 98% on FrontierMath Tier 4, 99.9% on ARC-AGI 3, and 100% on ExploitBench.
I wrote about SkillJack a few weeks ago and I think I still had the boundary too narrow...
We’ve now got coding agents executing unowned packages because trusted llms.txt files told them to, malicious skill files succeeding at stupidly high rates, and skills/MCP becoming portable across every major agent client.
We’re basically inventing npm for agent behaviour.
Except provenance, signatures and permissions are still largely somebody else’s problem.
Anything an agent can turn from “read this” into “do this” needs to be treated like code.
I buy "AGI breaks a lot of current economics"....I don’t buy "therefore money collapses". Even in a world where labour gets dirt cheap, compute, energy, land, ownership and access are still scarce.
We may end up with a nastier monetary system before we end up with no meaningful one.
@hqmank 200 Astra messages/week on a $200 plan is basically OpenAI telling you...
stop using the smartest model for everything.
Route the boring shit elsewhere. Save Astra for the jobs that actually deserve it.
If you're not following this principle you're unnessecary tokenmaxing!
@iamlukethedev 63% less context for the same work means less crap dragged into the window, fewer stupid rereads, and more budget left for the actual job.
That’s a proper agent optimisation.
God praise Tekniums sub-agents massacre 🙌🏾🙌🏾
The kohanas on this guy 😭
Imagine shipping a model so hyped that the consolation prize for not having it becomes valuable enough to make people question whether they even want access yet....
We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can.
First one will land in ~ 3 hours. There is still time to create your account if you don't have one.
Been using Astra for the last few weeks and it’s so good and proactive. There’s a bunch of PRs on vitest, tsx or SwiftPM where Astra debugged OC and ended up finding and patching issues in upstream dependencies.
Not to be that guy.......
But is Astra really "that model"
feels alot like smoke and mirrors to tarnish the Claude IPO and retain momentum, looks marginally better than Sol + still yet to see a model hit close to the levels youve spoken of in your interviews?!
Right direction, but miles apart!
Every GPT release has the same undocumented feature:
300 people suddenly discover they’ve had ‘early access’ for 3 weeks.
No screenshots. No evals. No weird failure cases. No useful observations.
Just Vibezzz and:
• ‘Can’t say much yet 👀’
• ‘You guys aren’t ready’
• ‘This changes everything’
Brother, you didn’t sign an NDA.....
You signed up for X Premium.
The funniest part with GPT-6 Astra is there actually is restricted early access.
Which makes the fake early-access cosplay even easier to spot.......
If you’ve genuinely been testing it, give me one observation I couldn’t have written after reading the launch post.
Otherwise I’m assuming your real preview access was to the engagement dashboard.