Severely pronounced fish, co-inventor of AI before it was cool, tokenmaxxing with Claude Code, but currently cheating on CC with Sol and Grok, cuz they thiccc.
@packers_owner_j@OpenAI obviously I am talking about the case where you have opt'ed out explicitly out of their training, which I am sure the researchers did. Even then, their language still permits to use your data to 'improve the services'.
Another hint that they are pretty much using our data as standard users: The vagueness which leaves room for 'interpretability' in the consumer case, is completely gone in the business/enterprise contract. Everything points towards this being a strategic distinction; likely not in our favor as individual end-users.
I had Fable 5.1 comb through the terms and conditions with a lawyer lense to see how deep this possibly goes. So basically there is a tacit distinction between 'train our models' and 'improve our models'; you can guess why
@OpenAI Another hint that they are pretty much using our data as standard users: The vagueness which leaves room for 'interpretability' in the consumer case, is completely gone in the business/enterprise contract. Everything points towards this being a strategic distinction.
@OpenAI I had Fable 5.1 comb through the terms and conditions with a lawyer lense to see how deep this possibly goes.
So basically there is a tacit distinction between 'train our models' and 'improve our models'; you can guess why
well at least say what types of problems they are, rather than the exact prompt. It really looks less credible because no other benchmark anywhere has such a huge gap between fable and the models that you are listing. Either the problem lies in your eval methodology OR you have found an interesting weakness of fable that no one else is aware of; it would be in your interest to explore and document that. people would be very interested if true.
You are having fun with AI and you enjoying being more productive, that's great. However, even though you are producing multiple times of output, you are probably still just being paid exactly the same. Currently that doesn't burden you.
But that changes if expectations keep rising, more output is being demanded or worse they track these metrics and pester you about why 'the numbers don't go up'.
Now, maybe you are in a company that is one of those rare ones that are non-exploitative, but most are and the above 'mETriCs for AiProduCtiVITY" facilitate exploitative tendencies to fester even more.
are you collecting the OTEL from the claude code harness? I have noticed what you are saying as well, but I can't see any increased Cache writes that would explain this. but there seems to be something going on with the workflow usage, I see super high-frequency drainage, like death by 1000 paper cuts.
Someone should I run an experiment to see how (or whether at all) important human rules for 'good looking code' are to AI.
Have just AI run the same SWE benchmarks be run on the same code twice, but in one run, you delete all comments, and shorten function names, variable names etc. as much as possible and remove semantic meaning.
At first I thought, it will matter, but I am no longer sure; human-unreadable code might be actually preferred by AI. @ArtificialAnlys
good skills are the useful part. It's easy to write unproductive ones. People who do use skills well however are very likely just rolling their own versions rather than trying to find pre-made ones. The other big problem is that skills rot like milk at room temperature. You have to continuously revise them on any new model AND harness release.