Really cool. nvidia based open model coalition
Nvidia offers DGX Cloud to AI foundation model labs for new Nemotron Coalition - DCD https://t.co/NX9VC5YBqD
@eqhylxx@vllm_project@nvidia Hi @eqhylxx thank you for sharing, is the nvidia vLLM tutorial sufficient to rerun your experiments on dgx spark? I was going to try this on mine. thank you.
I was able to try @yutori_ai as a beta user. collecting latest financial trends to gathering neurips after events from an online agent. I suggest folks to try!
A few months ago, I started a series of blog posts to explain what it is like to work as a researcher in the industry in the hope of documenting some useful information for those facing this potential transition (especially junior folks). Today, I've published the last part
🧵
Last week our paper on SynthID-text, a method from @GoogleDeepMind for watermarking LLM-generated text, was published in @Nature. 🧵
https://t.co/5IUKXkPjs7
AI companies are constructing private highways on public land—commercializing closed products built on public data, public resources, open research, and the open web.
We can’t rely on companies to build everything we need from AI, and we can’t afford the risk that they won’t.
New short course on Pretraining LLMs! Developed with @UpstageAI and taught by their CEO @hunkims and CSO @echojuliett.
While prompting or fine-tuning existing models works well for many general language tasks, pretraining is valuable for specialized domains or languages with limited representation in current models.
This course walks you through the LLM pretraining pipeline:
1. Data preparation: Learn to source, clean, and prepare training data using HuggingFace.
2. Model architecture: Configure transformer networks, including modifying existing models.
3. Training: Set up and run training using open-source libraries.
4. Evaluation: Benchmark performance using popular evaluation strategies.
As an example use case, you'll also compare the output of a base model with its fine-tuned and further pretrained variants, to see the impact of pretraining on a model's ability to write Python.
The course also explores an innovative technique called depth up-scaling, which Upstage used to train their Solar model family, reducing pretraining compute costs by up to 70%. This technique works by first duplicating layers of a smaller pretrained model to form a larger model, and then further pretraining the result.
Sign up here! https://t.co/IjYmNPR7sd
Just published the second part of my blog post to reflect on my journey as an early-career industry researcher, where I talk about how I would approach choosing a team today and what I'd factor into my consideration.
Post: https://t.co/SCGeYkV1D9