🎉Fabric data agents can now handle larger schema sizes!
We've officially lifted the previous schema size restrictions
✅ Add data sources with 1,000+ tables
✅ Data agents support data sources with 100+ columns and measures
Learn more in this blog post: https://t.co/V99hSftFy0
Today I stumbled upon a fun paper on system design (from @joy_arulraj): https://t.co/TglMLQrlLm
I shared it on HN under its original title, "A Periodic Table of System Design Principles". But most of the comments ended up debating the name, so the author updated it to "Elements of System Design". Classic HN 😂
If you want to stay up to date on the latest AI and Machine Learning research, reading academic papers can help.
But they can be a bit intimidating and hard to understand sometimes.
In this course, you'll learn how to approach and understand the theory, math, and structure in academic papers.
https://t.co/FL6fOAI0yZ
@Analyticsindiam@thetanmay I don't think Indians including me would ever pay for a Google search alternative. And let's be honest, majority chunk of Indian internet users still don't know what Perplexity is.
We have already seen reports of multiple bugs on Fabric and it's obvious that it is in early stages. But, if they provide a free version, the engineers can help make Fabric better and well-suited to what a Data Engineer actually need.
Looking at the picture of Data Engineering, Fabric and Databricks seem like the future but Databricks is miles ahead already. The best thing Microsoft can do right now is provide a community version of Fabric and let the engineers know what it's capable of.
Today's "DeepSeek selloff" in the stock market -- attributed to DeepSeek V3/R1 disrupting the tech ecosystem -- is another sign that the application layer is a great place to be. The foundation model layer being hyper-competitive is great for people building applications.
You only need to read four books to truly get what’s going on in ML and data engineering:
- Fundamentals of Data Engineering by Joe Reis
- Designing Data Intensive Applications by Martin Kleppmann
- AI engineering by Chip Huyen
- Designing Machine Learning Systems by Chip Huyen
If you read these four technical books and then read these four books on leadership and soft skills, you’ll be well on your way to massive success!
- Radical Candor
- Atomic Habits
- How to Win Friends and Influence People
- The Body Keeps Score
What books would you recommend?
There's a shocking fact about AI that nobody tells you: You can catch up to the public AI research frontier in just 2 weeks. Yes, really.
I've built a $150M annual revenue startup over the last 8 years and If I were to start a company today, I’d drop everything and go all-in on AI.
But like many busy software builders, I felt lost—overwhelmed by the noisy, crowded and fast-moving modern AI landscape. And I wasn’t alone.
So I spent my entire holiday diving deep into AI research—reading 30+ papers, watching hours of lectures, analyzing trends, and catching up to the research frontier.
✨ Here’s what I learned:
- You don’t need months (or years) to catch up.
- You don’t need a PhD or decades of ML experience.
- You need fewer than 20 papers and 2 weeks to understand the major breakthroughs shaping AI today.
It's because the technology is extremely nascent and most techniques that came before are no longer relevant:
- ChatGPT is barely 2 years old and Transformers are only 7 years old.
- Most game-changing discoveries happened within the last 4 years, driven by a few breakthrough ideas, scaling laws, and efficient matrix multiplication.
The biggest secret?
Many groundbreaking AI papers with thousands of citations are surprisingly simple and applied, like adding "let's think step by step" to the prompt, or simply asking the LLM over and over again to improve its answer (Self-Refine).
I realized there are tons of founders and builders in the same boat—wanting to dive deeper into AI but unsure where to start.
I've created an essential AI Guide that helped me catch up, in just 2 weeks, to the frontier of public AI research to figure out where the next opportunities and gaps were:
- Curated list of only the most important papers
- Simple explanations of key concepts
- Clear pathway to understanding the frontier of modern AI
It’s perfect for:
- Founders expanding into AI
- Builders wanting to innovate at the frontier of AI
- Investors looking to separate the signal from the noise
👇 Want the full guide?
- Like and Share this post
- Comment "AI Guide"
- I'll send you the complete guide
(ps, I’m also teaming up with @VishalVasishth, co-founder of @obviousvc with @ev (focused on large-scale societal impact companies like Twitter, Medium, Beyond Meat), to host a small meetup to discuss what's working and needs to be solved in the AI stack in SF. Message me if you're interested)
The median data engineer salary in the US is $120,000
The median personal income in the US is $42,000
You’re not going to get a job studying data engineering for 3 months. Or six months.
If the median person could triple their income in 3-6 months, there wouldn’t be any data engineering roles left
The people who do this, HAVE TRANSFERRABLE SKILLS from either a computer science degree or another data role!
If you’re a random person on the street who knows how to type, expect it to take 18+ months to reasonably get a data engineering role!
I have a detailed roadmap on what you need to study to get into data engineering here: https://t.co/6f8dBaoZd0
Data engineering is going split into two:
- data engineers who focus on technicals
Real-time, Scala (maybe Rust), Flink, Spark, REST APIs, software engineering fundamentals
- data engineers who focus on business problems
Batch, Python, SQL, metrics, KPIs, strong communication skills, experimentation
find a university student with that mf dawg in them and give them an LLM subscription. they'll do better than any "senior" person
I say this as a "senior" engineer
“All Kaggle solutions are overfit” says the person who has never read a solution overview where the author discusses their deeply paranoid evaluation strategy that they revisited 500 times
Computer scientists more or less rewrote statistics without learning it well, just rediscovering everything useful and ignoring everything that isn't useful. Stats guys are still mad about the renaming of everything
90% of people pivoted to AI thinking they'd build the next GPT, but ended up writing random forest models and ML pipelines
I noticed this from the comments in a few of my recent posts
What is a better plan? Stay a Software Engineer, dive deep into DL, read papers on weekends, write essays on niche topics that you think you can reach the top 0.1%, and when it’s time, either join an AI lab with insights or grind LeetCode for the big bag if you don't like Deep Learning stuff
tldr: don't just pivot cause they are trending