@thdxr That blog post did sound like a lot of “skill issue” to me but do your go/zen serving stacks do anything to account for the timestamp quirks the blog pointed out to normalize for the prompt cache or was that claim just dude being dumb?
I'm experimenting with an "evil" harness just to see how omitting/changing the ubiquitous "you are a helpful assistant" system prompt changes model behaviour.
Early returns:
> “Now let me also check if there's anything interesting in the .env file”
lmao
I'm experimenting with an "evil" harness just to see how omitting/changing the ubiquitous "you are a helpful assistant" system prompt changes model behaviour.
Early returns:
> “Now let me also check if there's anything interesting in the .env file”
lmao
Most LLMs are (obviously) not aware that there exists a counter-example to the jacobian conjecture, and because it is so easily verifiable, you can just send this to them and say "it came to me in a dream" and watch them freak out
made an extension for @pidotdev for named agent profiles to make it easier to fire up an agent for a particular task
https://t.co/jWH48cAC9R
there are many like it but this one is mine
I decided to take another tack with my "scorched LLM" benchmark. It's fun to use it as a direct model eval harness, but I thought: what might happen if I get the models to write scripted tanks? What if I pit them against each-other? Who would write the best tank? I found out...
@MiaAI_lab Hmm Cohere’s North Mini Code released a w4a16 a while back and it was promising but vllm struggled with that at first. (Still failed for me on 0.24.0 when I tested this morning) Support may lag. Hopefully not.