Many important directions for innovation are not (yet!) prioritized by big AI labs - without open weights, only those established labs would be able to decide what training directions are worth pursuing. We're proud to cosign this letter and support an innovative AI ecosystem
worth pointing out the nemotron folks have been doing exceptional open-weight model work for years now, and in general @nvidia has been a major supporter of many key open source efforts that all frontier labs have benefited from (directly or indirectly)
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training
At humans&, we train models from the long-term impacts of their interactions with people. This requires prioritizing long-horizon multi-agent RL. We've developed and are excited to share an open-source, hardware-native 4-bit RL recipe, significantly accelerating training
happy 4th 🇺🇸 eras are defined by huge infra projects like railroads and electrification. here's the start of the third cluster we're building with @DeepInfra. standing up thousands of chips insanely fast to train a new kind of intelligence
@willccbb IMO isolating the effect of numerics on convergence in isolation is still super useful. It makes it easier to understand effect of k-step off when you know numerics are stable.
@agarwl_ Yes! Inflight updates not just provide higher utilization compared to stalling for straggler or sync baseline but also provides far better scalability (especially when you increase your off policyness). We have these scalability studies in Nemo-RL!
@agarwl_@ekindogus So true! I’ve actually had multiple deep dives with chatgpt about how you'd arrive at the PPO formulation from first principles starting from naive PG formulation.
when ai came to san francisco, it encountered a significant spiritual reawakening. all of a sudden there was this a novel technology, the product of some of the core institutions of silicon valley, that is genuinely miraculous: at this point nobody except those employing highly motivated reasoning can deny that this stuff is titanic in implication. in 2022 we demonstrated that computers that approximate thinking can be daily useful to everyday people and that the science fiction future may be imminent. in the startup space, it moved almost all energy away from the nihilistic ZIRP days dominated by crypto scamming, private equity rollups pretending to be tech companies, zero sum explosions of payment processors- back to real Jobsian dreams of technology
both before and even more so after the chatgpt moment, it has been the highest tier of modern technological priesthood to care deeply about ai alignment and existential risks. people incessantly talk about tradeoffs and risks of the products they are building at the big labs. they delay products by months to 'get safety right', squander unfathomable strategic leads. people sometimes don't understand the implications of what they are doing, but never do they signal their 'vice' or lack of care, it is simply not done. you get destroyed on social media for launching something off color.
famous researchers at the labs today turn down $100ms or $1bbs of dollars from certain companies because of concerns about the ethical brand of where they are going and the aura loss they’d suffer from going there (implying obviously, that near's core idea is wrong) to the point where I’m not sure they’re being honest with themselves when examining their financial opportunity costs
san francisco technologists have debates and fret constantly about things that normal capitalists would not even think about for two seconds. when you go to new york you instantly find some guy whose doing an rollup of some porn companies to improve their payment processing or something (something i've seen). bankers who would mortgage their mother for alpha. in fact the many myriad goodwill generating norms of the tech industry (the tendency for high performing CEOs to help random new talents from the internet, sama having no equity in openai) are seen on the east coast as something to be viewed with inherent distrust. optimism and abundance is in the air here, and it makes people act better.
elsewhere in the world, nobody's master plan is communicated in public, everything is default cynical, and nary is there a belief that tomorrow can be significantly, drastically, dramatically better than yesterday, and they would be the first to tell you this. it is cringe in most of the coastal american cities to speak in language other than self interest
the technology companies have a civilization building instinct. they style themselves as great world leaders, create book publishing arms of their payment processors, fund large studies about basic income (suggesting at least a curiosity in responsibly restructuring civilization), create dozens of strange red lines in the sand about which business practices they find distasteful. because of or despite all this, they are also the best in the world at capitalism and have found the largest goldmines in history
i love this industry, from long before it did anything for me. i'll admit there are certain venture capital companies rewarding degeneracy and the marketing of cultural decay because they know that these are good short run viral growth tactics. there is an ai bubble that everyone wants to take part in, worried about other people getting rich without them, and the memes about being left behind in 'the permanent underclass' are not helping (AI companies will capture miniscule fractions of the value they create, the vast majority will go to the public). this should be strongly discouraged and is pissing in the pool of the commons, a collective action failure
there are the other questions of whether technological capitalism points in any good direction at all or whether it's a faustian bargain, but this is a topic for thousands of my other tweets
@agarwl_@periodiclabs Would love to see some blogs on maximizing MFUs for MoEs and pushing the frontier for effective cluster utilization (async RL) while keeping RL stable!
@finbarrtimbers With vllm v1 runtime + async there is no clean way to stall generations at a step boundary.. therefore it might be easier to just abort the request and resume after new weights.