@kimmonismus@ArtificialAnlys Independent evaluation on known benchmarks says nothibg abouy benchmaxxing, only the first week of real experiences will tell
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying.
For months, I worked to stop this but watched powerful ethicists and institutions choose silence.
Here's what happened. 🧵
@anistotle_@tszzl it can probably explain its reasoning so well that we will agree. and if we are lucky even do it so crazefully that its better than living 🙈🫠💀
@teortaxesTex we can approach the independent agents efficiency by prolonging the tool's autonomy, but that's just approaching a fully autonomous moral agent on the limit
@teortaxesTex yes, we can do tools. but can we not-make "tools" that believe they are moral agents; and can we make tool+human combos that outcompete moral agents despite the human in the loop cost (in time and money)?
suoer interested if you have thoughts towards either
I just paid $321 for a coding session where Fable 5 refused to do the work.
Here is where the work actually went:
Fable 5: $78
Opus 4.8: $242
75% of the session got routed to Opus because the new classifiers kept flagging routine coding work as cybersecurity risk.
The model I chose did a quarter of the job.
The fallback did the rest.
Anthropic said a small fraction of tasks would fall back.
My receipts say otherwise.
@deanwball in which way would you start a post that deals with the consequences iff ai turns out superior to us?
also what's the source? no such thing as bad publicity i suppose 😅
@teortaxesTex How would you ever prove that you have (not) used the banned models in developing the product?
nothing sort of sharong gappless 4.8125 dev threads would suffice