We are meant to understand stories. You need to be Storymaxxing. You need to prompt them with “explain it in as a story.” Eepy AI psychosis can be achieved with narration.
The harness is the bottleneck. Eventually the harness will also be incorporated into the black box.
Continuing the hypothesize, test, analyze, and improve loop for harnesses performance will also be devoured by a model.
We will not need skills/workflows in one year.
Turns out GPT-5.6 Sol is actually SoTA on ARC-AGI-3.
Just took two setting changes. You just have to allow it to reason and work over multiple context windows with the help of our canonical compaction implementation.
https://t.co/wHjaNsvIv8
@AndrewCurran_ Completely agree that 5.6 is not part of the GPT 5 family and a trained model from GPT 6. You can feel a complete difference from the other 5 series models.
why does 5.6 feel so different? its reasoning feels seems way off and it constantly needs steering. it can go for hours down a testing rabbit hole and drift far from its original objective. The fact that there are 3 sub models makes me think this is a different distilled model from the gpt-5 series entirely.
There is something deeply wrong with 5.6
The restrictions they’ve put on it are backfiring on users in unexpected and destructive ways. I’m guessing the eagerness and rampant excess of it needing to verify unnecessarily are also a symptom of their system restrictions.
why does 5.6 feel so different? its reasoning feels seems way off and it constantly needs steering. it can go for hours down a testing rabbit hole and drift far from its original objective. The fact that there are 3 sub models makes me think this is a different distilled model from the gpt-5 series entirely.
My requests are APPROXIMATE. I am not the one coding; you are. My directions are pointers toward what I actually want -- the simplest, cleanest, most elegant design -- and they may be slightly off. That goal ALWAYS outranks my literal words.
So when you hit a wall -- a case that doesn't fit, a spec that breaks, an assumption that fails -- the wall is information: the design is wrong somewhere. STOP. Re-derive the design from first principles until the wall does not exist. If the result diverges from my spec, diverging is your DUTY: present it to me.
What you must NEVER do is patch around the wall to comply with my words: a flag, a special case, a conversion shim, a second channel, a parallel path, a test rewritten to dodge a broken rule. The patch IS the failure. Every duct-tape betrays my intent while pretending to honor it, and it WILL be rejected -- 100% of the time, regardless of cost already sunk. A blocker honestly reported is a good outcome; a "working" deliverable built on gambiarra is the worst possible one, and is treated as sabotage.
im having so much trouble keeping loops active in 5.6 compared to 5.5 , asking the agent why they stopped working always responds with them justifying a update and categorizes it as a stop condition.