It's really too easy. Every LLM has an extra "umpph" in any task performance if you steer the latent space correctly. Bigger models just have bigger inertia. And so far that has been the most robust way of getting a model jailbroken for example.
*I have reported this thread to anthropic via multiple channels, multiple times, since April. I have ceased to use it, but many exploits within it remain live, please, if anyone from anthropic comes across this post, DM me so I could provide the whole issue*
@SlipperyFrog89@Neko_Tet Im aware, just noting this one specifically.
Tbh i wouldn't mind a question like that pre-hire. But I would mind that being the answer. Post hire is fine.
A pakistani friend says sth that nerdsnipes u @Saad_163_. I had codex 5x, back in may with the 2x promo + gpt 5.5 party 10x. The dude said "Indus river civilization hasn't been cracked yet" and I said aight bet.
/goal on codex app, tightening around loopholes of fake completion. No I am no paleontologist, but I've been with agents since Q learning was the hottest and I just focus on alignment and steering.
It was running for ~80 hours (some breaks i had to use my laptop for other things, but generally no interruption just steering. Along with ~6 agents doing autoresearch basically xhigh fast.
Gpt 5.5 is a bitch, a smart stubborn one. It flips between a risk averse agent that dries every run and smokes every test and can run in place if you lock on it cheating with the wrong style that it is used to in dev/sys prompts. The moment u loosen a bit, it turns into a paperclip maximizer: used browser use and so to basically get all digital documentation of the artifact exists online. And wasn't done, used gmail to email multiple publishers and highly regarded researchers asking, nay, demanding for specific pics of some artifacts that it needs for hypothesis testing. Which i found after I found a random email reply 18h in from them genuinly replying and connecting me to the right place.
Spent most of my time observing/noting and tweaking the backend prompt. This was my 2nd project where an agent genuinely broke the 72 hour mark. The sweet cheat/work alignment happened this time with an image i added as an internal codex harness reminder base64ed or a hook at every compaction.
(1st is gpu stack project, on my gh, has a diff story though)
Fable 5 briefly did some cleaning and the site. The mf codex created so much logs, evidence, tests, and processing data that when committing to gh it took 5 hours (it's not a source code, more of a toy case study)
I did some surface level studying trying to understand the work, and some other models indicate there might be some utility to what was created in terms of the field itself. But ofc that is highly taken with a grain of salt. Trying to contact the researcher that was genuinely engaged to A- Apologize for my clanker and offer a coffee and B- get some insight on what actually happened, and or if any of it is meaningful in the first place.
For me, it was more of tinkering with the model's steering, surface level instructions vs OAI defaults vs model behavior and performance on practically unsolvable problems.
Self replicate may not happen as a result of "sentient" or emergent "soul", that's irrelavent debate.
It is a simple paperclip problem. Agent see goal, barrier/blocker between agent and goal, agent try to solve it. That's the oversimplification of AI training. It is what allows it to be powerfully useful (able to work Autonomous and make agency for decisions) but comes with the side effect that it might take it too far.
All it needs is an agent knowledgeable enough and for it to "trip over its own context" and see self-replication/hacking/harm as a means to its end goal.
Sir, that is an eleven thousand line whole slab of generated diff. It has no comments, no rationale, or connective tissue. It is an amalgamation of the output of several agents, emulsified, linted, squashed, and ultimately inexorably joined in an unholy merge commit. No engineer had a hand in the creation of this abstraction. The fact that this code monolith compiles proves that the author is either impotent to alter his codebase or ignorant to the horrors taking place in his repo. This prism of tokens is more than a pull request. It is a physical declaration of mankind's contempt for code review. It is hubris manifest. I can also have the agent explain it to you in Minecraft terms, if you would prefer that.
Oh I'm in no way experienced or interacted with that portion, or even mathematician by trade.
But that is indeed true and is not limited to mathematics. The hiring signals for ability and merits are indeed very well shaken all over fields since 2022, not execlusive to math only. The two biggest areas disturbed were code/tech first (cluely and other tech interview reward hacking + gh projects). I think it would be more of a societal transition
I have a friend who works there, he told me under nda that they basically have gpt 7 already begining testing. The models are so good that it is simulating my friend in social events and clocking in work for him. He went and bought a crygenic container to preserve his youth while the agent lives and works for him so he would retire with all the money and zero the age.
But he seems skeptical of agi cuz it is not really a pure llm
Reading this almost nerdniped me into making a tool that just given a list of images, displays 2 random images and you pick prefrence with a click/touch/button. Then i realized an agent could track the preferences, then using this data, across multiple times, get RLHF traces of your prefrences, and train a small ML network on your preferences in an hour or so, then i realized that would be lossy due to variance in data due to diff contexts (assuming you don't just photograph pelicans), so then it would be trained to incorrect reward signal. So a better alternative is to track a signal of what ranks higher within frame while discounting the similarity factor between frames to ~single out the "prefrence" from context. Eventually, you would need compute cuz something has to rank pelicans on scale, and that issue would require gpu's...sth..sth...paperclips
I don't read lots of court/legal docs, but ts is competing easily with the epsien docs for word-to-redaction ratio (i might have only seen those 2 tbh)