Suppose the AI field is spending 20x as much on capabilities as safety (probably an underestimate). Shifting 10% of R&D to safety would 3x safety. We should absolutely be building a mechanism to "pace" progress, but meanwhile, increasing safety spending is low-hanging fruit (and, if you're an accelerator: this would reduce incidents that trigger pushback / regulation).
See e.g. tweets from @yonashav (https://t.co/u7icqpCmld) and @fleetingbytes (https://t.co/CIlLpX4Uvv)
@datagenproc@Raemon777@ohabryka Ask an LLM to judge the value of a codebase and filter? I don't really have good ideas. Low value and low quality code can be embedded within largely high value projects unless the project has standards not only for quality, but also for value
@Raemon777@datagenproc@ohabryka I agree. I'd guess that there's way more low-value code, because before, code was mostly written only if you had a good economic reason, whereas now anyone can create code without a good reason.
@datagenproc@ohabryka@Raemon777 Further data is very annoying to get from the github api, but my guess is commits and PRs have increased at the same rate and LoC even faster.
@Bayesian0_0@gwern If you measure across a wide enough x range, and your noise doesn't grow proportionally, the r^2 will be very high. Effective compute varies by orders of magnitude so there is no surprise it has high r^2.
@YafahEdelman Do we know whether it tends to be a sigmoid or some other function like rescaled arctan? There are so many benchmarks that we should have enough data by now.
@Bayesian0_0 Take care to account for label noise. Many benchmarks are poor quality and have say 10% broken questions, so if the human baseline is 85% and the true ceiling is 90%, the benchmark could be mostly measuring the AI's ability to do broken questions.
@ben_j_todd If I had more time I would do an updated version of this where we found similar time horizon slopes in easy and hard domains, but with hard-to-verify tasks. https://t.co/ei9lQwfC3a
@ChrisPainterYup@testingham Everyone needs to be familiar with some of these. It is already the case that agents make ~unlimited progress on some tasks (log-linear or better) when you spend enough, so you need returns to expenditure to say anything meaningful.
@YafahEdelman This idea likely gets better as inference scaling improves, the uneconomical region grows, and the perf vs tokens curve gets smoother. The slope of the curve also tells us how close they are to human inference scaling.
@eli_lifland My guess:
- There are 5+ industries of similar sizes, due to complementarity. Oil, tech, health are all necessary for the economy
- Amazon+Walmart are similar due to luck and/or antitrust. If one were >2x the other (~1T revenue), it would be 2x larger than all other companies
@ptrschmdtnlsn@seventhmeal If this had happened it could have accelerated the mathematical understanding of probability by centuries, since it historically came from gambling.
@simple_as_grass these are *certified* random numbers, individually inspected by a panel of experts to guarantee randomness. This process likely cost £10,000+ and is far superior to most of the public literature, which uses statistical tests and so can't provide hard guarantees
@emollick These are $500B+ companies so if they could just scale version number they would have. IMHO Anthropic and GDM are completely version number bottlenecked, and if GDM could copy whatever version magic OpenAI has into Gemini and get to 5+ themselves, they'd easily be in the lead.
@SpencrGreenberg Maybe this is because we remember at 2x speed. Speaking words for the first time happens at 1x speed, but when we're just remember them there's less mental work and so 1x feels uncomfortably slow?
@Scholars_Stage The US has dropped 2000+ bombs on Iran. With ASI, each bomb could be replaced with >1000 drones that can use their perfect coordination to defeat >1 soldier each without collateral damage. Superhuman negotiation would also end the war sooner, on terms more favorable to the US.
@Robotbeat It helps that Merlin 1D+ has very high TWR. Falcon 9 v1.0 (Merlin 1C) had ~3.1% payload fraction, now it's over 4.1%. Some of the engine performance required subcooling.
Other factors: no solids, large size, other rockets are optimized for minimizing use of expensive parts.
@LinkofSunshine God the economist would draw a random variable X_i for each soul, only known to Himself. If person i has > X_i unrepented sins they are damned. This incentivizes marginal acts of repentance while maintaining God's omniscience.
@aakashgupta Zuck is ~1M times wealthier than the average American, so this feels to him like leaving a lawnmower on while it uses 2.8 cups (0.6768 liters) of gas.