New #Dodgers draftee Gavin Van Kempen was having an excellent season for East Carolina before he got hurt in March
Through six starts, he posted a 1.84 ERA, 39% strikeout rate and 8% walk rate. After suffering an arm injury in March, Van Kempen had surgery in May
🤗🤗🤗introducing Hugging Science -- the home of AI for science 🤗🤗🤗
open models and datasets are the powerhouse of science (see the PDB), but finding the models and data you actually need for your breakthrough is hard af
you shouldn't need to scrape arxiv, own your own wetlab, fight a custom HDF5 parser, build a fusion stellarator, and beg for compute before you've trained a single epoch
so we're changing that
we've put all the best science on @huggingface in one place:
- 78GB of genomics data
- 11TB of PDE simulations
- 100M cell profiles
- 9T DNA base pairs
- 13M molecular trajectories
- 400k medical QA pairs
and much more, all open, and all ready for training (+ you can also now filter and search by domain, task, and keyword)
we've put together all the biggest releases from our partners at NASA, Google, OpenAI, Meta FAIR, Arc Institute, Ginkgo, SandboxAQ, Proxima Fusion, NVIDIA, Ai2, OpenADMET, InstaDeep, Future House, Polymathic AI, LeMaterial, Earth Species Project, Merck, and Eve Bio
if you're not sure where you fit in -- work on open challenges for problems that matter: including fusion stellarator design, ADMET, antibody developability, multilingual medicine, catalysis and materials, and scientific reasoning.
we're already changing how science gets done:
a fusion startup needed a benchmark for stellarator plasma confinement that didn't exist. @proximafusion shipped ConStellaration on Hugging Science: a leaderboard, dataset, and eval metrics, all in one place.
a drug discovery team wanted to predict hPXR induction. OpenADMET put up a blind challenge: 11,000+ compounds assayed at Octant, 513 held out, two tracks (pEC50 + structure). Anyone in the world can train and submit.
an antibody team at @Ginkgo released GDPa1, a developability dataset for stability, manufacturability, and immunogenicity prediction, with a live leaderboard scoring every submission.
if you know a problem the ML community should be working on, let us know. make a challenge! this is about putting all the tools for solving science in one place. so we can hillclimb!
→ https://t.co/T4l4r1lDz0
@iamtrask Yes, you've got it right from my point of view - regulated industries absolutely benefit from Local AI setups.
Outside of simple use cases where you have to serve many users and models, this blog post is a great reference for an enterprise stack
https://t.co/Wuik4p0Xbk
Turns out with claude code, my decades long strategy of NOT deeply learning:
- regexs
- sql
- nginx confs
- elaborate shell commands
- advanced shell scripting
- any javascript framework
- perf optimization
- webpack, cdns, bundlers
- 1000 other things
...was entirely correct.
Sean McDermott is one of the best leaders of men I’ve been around. From ending the drought to leading through adversity, he galvanized this community. No doubt he will continue to win and positively impact lives wherever God places him next.
IBM dropped CUGA, open-source enterprise agent to automate boring tasks 🔥
> given workspace files, it writes and executes code to accomplish any task 🤯
> comes with a ton of tools built for enterprise tasks, supports MCPs
> plug in your favorite LLM 👏
here's a small demo where it retrieves info from a file, calculates revenue by writing code, and drafts an e-mail 🤯
they release code, a blog and a demo 🙌🏻 you can run this locally
Training LLMs end to end is hard. Very excited to share our new blog (book?) that cover the full pipeline: pre-training, post-training and infra. 200+ pages of what worked, what didn’t, and how to make it run reliably
https://t.co/iN2JtWhn23
In order to celebrate the release of the print version for the Ultra-Scale Playbook (of which I have no affiliation with and love deeply), I'm going to be giving away 5 copies!
To enter, simply like + retweet this tweet. Winners will be selected at random 10AM EST on the 13th
I think Confabulation captures the behavior better: Confabulation is a memory error where a person unintentionally recalls false or distorted memories, and believes them to be accurate.
w OpenAI adding a router in GPT-5 its a good time to say that one of the ways open models win is by routers being easy to train and then they can select between 1000s of specialized models that no one company could train on their own in order to make fun model networks.
We've been working on extracting information in templatized PDFs for the last couple of years, leveraging the best of LLMs and classical data extraction techniques.
Our latest technique, TWIX, has the best of all worlds: beats Azure DI, AWS Textract, or LLM-based approaches by over 25% in precision AND recall, and is multiple orders of magnitude cheaper and faster. TWIX is now open-source, with an API plus an optional front-end for you to tweak the extraction results.
Try it out on your gnarliest datasets and let us know what you think!
Led by Yiming Lin, w/ Yash Jain, Mawil Hasan, @alvinkcheung