Harvard just open-sourced its entire ML Systems curriculum.
Free. Public. 6 pillars. Hundreds of pages.
And it won't get most data scientists any closer to a $150K AI role.
Here's why.
i'm actually more interested in popular history works that historians universally agree "yeah, that's a solid one."
first one that comes to mind is "1491: New Revelations of the Americas Before Columbus" by @CharlesCMann
I Wrote a New Book!!!
Optimization: A Bootcamp for Machine Learning, Inverse Problems, and Control
Pre-Order Now (July 31)
https://t.co/EoDMFapUUf
Coming Soon:
* Free PDF on website
* YouTube Videos for entire book
* Python code on GitHub
We're very lucky that so much good learning material is publicly available on arxiv, such as this very complete 'Lectures on Differential Topology' by Riccardo Benedetti (400+ pages)
Covers subject such as bundles, transversality, bordism, Morse functions, Pontryagin-Thom construction, the basics of 4 manifolds ++
🔗👇👇👇👇
New essay: A Brief History of Bioinformatics Software
Although the word "bioinformatics" wasn't coined until 1970, the first computer program to analyze protein sequences, named COMPROTEIN, was published in 1962.
From our forthcoming book, "Making the Modern Laboratory."
should i publish this? basically i used old-style statistics (spectral analysis) to figure out the latent space of a bunch of recipes. Analyzed about 180k recipes and it lets you drill into cuisines and/or features. Also thinking of a seasoning directory to tell you what flavors pair well.
Harvard just open-sourced its entire ML Systems curriculum.
Free. Public. 6 pillars. Hundreds of pages.
And it won't get most data scientists any closer to a $150K AI role.
Here's why.
Harvard just made degrees worth $200k obsolete by open-sourcing its Senior AI Engineer roadmap
Stop paying for bootcamps. Prof. Vijay Janapa Reddi just put the entire ML Systems (CS249r) curriculum on GitHub.
If you master these 6 pillars, you're ahead of 99% of the field:
> Architecture
> Data Pipelines
> Production
> MLOps
> Edge AI
> Privacy
This is the "Black Box" of Big Tech infrastructure, open-sourced.
Read. Learn. Bookmark.
Book - https://t.co/997GB0s1fl
GitHub Repo -https://t.co/y8pDr4lGFt
[1/n]
Super excited to introduce PaperBanana 🍌! (PKU x Google Cloud AI)
As AI researchers, we often spend way too much time crafting diagrams and plots instead of focusing on the ideas 🤯. To rescue us from this burden, we built an Agentic Framework to auto-generate NeurIPS-quality paper illustrations!
📄 Paper: https://t.co/2NbQeEhzMv
🌐 Page: https://t.co/05dKkjVs7f
Key Features:
🌟 Human-like Workflow: Retrieve 🔍 -> Plan 📝 -> Style 🎨 -> Render 🖼️ -> Critique 🔄. This ensures both academic fidelity and aesthetics.
🌟 Versatile: Supports both illustrative diagrams and statistical plots.
🌟 Polishing: Also effective for polishing existing human-drawn diagrams.
Here are some example diagrams and plots generated by our PaperBanana:
What podcasts do you guys listen to regularly? I do Know Your Enemy, Interesting Times with Ross Douthat, Ones and Tooze, The Time of Monsters, some cinema podcasts, QAnon Anonymous, and In Bed with the Right
News, business, and current events:
- Odd Lots
- FT News Briefing
- Marketplace
More technical topics:
- Shift Key
- The Innovation and Diffusion Pod
- The Economic History Pod
- The Inequality pod
General research:
- Casual Inference
- Backstory
- New Books Network
Local:
- Brave Little State
- Vermont This Week
Understanding Deep Learning -- a 541-page PDF eBook and 68 coding exercises with Python code notebooks that cover all the topics in the book: https://t.co/a1vDJHSsUy
...or buy the book here: https://t.co/0fMrJf5yu7
Many asked how to so-called "break into data engineering". To be honest, if you just read whitepapers, you could go far.
I curated a list of specific data engineering papers.
- Lakehouse: Unify Data Warehousing and Advanced Analytics
- Ground: Data Context Service
- Spark: Cluster Computing with Working Sets
- Google File System
- Dataflow Model: Massive-Scale Data Processing
- MapReduce: Simplified Data Processing
- Dremel: Interactive Analysis of Web-Scale Datasets
- Data Mesh: Beyond Monolithic Data Lake
- MotherDuck: DuckDB in Cloud and Client
- Analytics Development Lifecycle (ADLC)
- Compiled and Vectorized Queries
More on https://t.co/t9B67hbU7R.