@hamandcheese Everything is verifiable but the constraint is more about the cost and delay of verification.
You can even “verify” a creative work like a novel by whether it’s a best seller.
But it’s not really viable to run millions of rollouts of that in a day for an RL agent to optimise.
@hamandcheese Everything is verifiable but the constraint is more about the cost and delay of verification.
You can even “verify” a creative work like a novel by whether it’s a best seller.
But it’s not really viable to run millions of rollouts of that in a day for an RL agent to optimise.
@FrancoisChauba1 Maybe you are vibe-coding an app about kevin costner or the population of paris?
Point being - do you not think that type of general knowldge is required or helpful to operate in the world effectively or complete tasks?
Nobody is saying that all software has to be open source. What they’re saying is that open-weight AI should be allowed. By implication, they’re rejecting your company’s incessant machinations to kneecap the open model ecosystem. Read the room.
No one should support an AI company as an investor or customer whose vision for our future is that we cede control to AI.
No one wants to live in this anti human future. No wonder so many people outside of tech have such negative views of AI.
https://t.co/uylIa7IUL0
There's no legal precedent that model outputs are IP. I'd argue it also doesn't stand up conceptually.
The frontier labs are under no obligation to have their APIs be available to these customers, and can protect model generations in other ways if they want to.
Many people seem to be arguing this is implausible because Fable hasn't been out for long. But Moonshot could have obtained unauthorized access before Fable was released via actors with prerelease access (Glasswing) or possibly by hacking Anthropic directly.
We doing gymnastics wtf. Theyre good enough to hack all the AI companies and access everything.
But also not good enough to maybe just be capable of training good ai.
Introducing Vals-Smith: turn your code base into a customized benchmark.
Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve.
New models ship every week. Vals-Smith tells you which one to trust with your code.
@blader I'm not sure they will eat skills.md as they specify the specific taste I have and the way I like a task to be completed.
Like when a new member joins a team, they dont just rip everything up and do things however they like, they onboard to the ways of working and build on top.
If US labs truly believed that the only reason Chinese open source is keeping up is because of distillation, they would have implemented KYC checks on customers to prevent this, and defend their trillion dollar valuations.
Wonder why we haven’t seen that yet…