@HyperTechInvest > It’s strong evidence that, so far, there are no signs of overbuilding and that, if anything, there is still a GPU shortage
We probably shouldn't look at spot prices to see signs of (future) oversupply; we should look at future pricing, if such a market/price even exists
Correct. No serious firm is likely to agree to sell their internal workflows/data so an AI lab can train on it and potentially improve the processes of a competitor. The ones who do are probably companies that are in need for cash, so optimizing their internal ops purely for Anthropic is .... a bit unlikely.
@echen It would be interesting to know what the bands of variance are on the SWE bench evals, ie: how many times has it been run before and after post-training :D
This train of logic assumes that demand/GPUs are uniform (ie: company A with its simpler email workloads requires the same level of 'frontier' capability or GPUs as company B) and that LLMs in the future will not be able to run on simpler or even consumer hardware.
Plus, such price setting is too simplistic - Company B might pay more for tokens because it's engaging in a higher risk & reward business, but what say if it were to fail/shutter - what happens to the compute contract and supply then?
@RadishHarmers These sort of time-boxed experiments are pretty difficult to execute. It's not easy to say how much of mathematical intelligence is purely derived from training on mathematics data as opposed to other pure sciences/logic/etc.
Plus the issue of data scale
@Just_another02@ArielKwiat Heavy arithmetic like that serve little purpose and is a pretty contrived task. But if you wanted to do 5.29392392 x 6.2929329 using a pure LLM - it will get it spot on.
@deepfates This is really a function of whether the sandbox/container allows for internet access (some data buyers don't, especially for sandboxes with GPU access) + whether there has to be a specification the vendor has to follow, ie: a docker-image with specific libraries bakes in.
@max_spero_ Turnitin outputs a lot of false positives unfortunately. Even multiple submissions of the same work can trigger flags. Turnitin scores are only useful with context and the exact details of the text that matches other work.
@multivitaman RLVR costs as much per row as RLHF data, often more in terms of human annotator spend.
RL env/agentic tasks are still human-written, so the grader is still just optimizing human-preference (ie the objective ground truth) but in a more systematic manner.