Confirming my second point. The "First Law" is absent in both courses and even the "Second Law" (d-separation) is nowhere to be seen. Coursera is hungry for a modern course in causal inference. Volunteers?
If you are having a hard time visualizing all the layers and matrix operations inside an LLM, then you'll love this website!
Check it out: https://t.co/5YTwn3XgdY
Correlation can be highly misleading!
Many solely rely on the correlation matrix to study the association between variables.
But unknown to them, the obtained statistic can be heavily driven by outliers.
This is evident from the image below.
The addition of just two outliers drastically changed:
- the correlation
- the regression fit
Thus, plotting the data is highly important.
This can save you from drawing wrong conclusions, which you may have drawn otherwise by solely looking at the summary statistics.
--
If you want to learn AI/ML engineering, I have put together a free PDF (530+ pages) with 150+ core DS/ML lessons.
Get here: https://t.co/M3Rh9zFc2M
--
👉 Over to you: What are some other measures you take when using summary statistics?
One of the mistakes you’ll make as a data engineer or data scientist early on in your career is not truly understanding the business requirements.
The business will come to you and ask for a real-time dashboard.
But they mean they want the data updated 3-4x a day or maybe they only look at the report once a week and at that moment the data should be as up-to-date as possible.
The business will ask for a machine learning model to help detect fraud.
But once you understand what the business generally knows is fraud, you'll realize you just need basic anomaly detection.
This isn't to say you don't ever need more ML models or need to deploy a near-real-time system.
But it's good to figure out what the business really needs before building what they describe.
If you'd like to learn more on the subject, here is a video on the topic!
How To Fast Track Your Data Engineering Career - Translating Business Requirements Into Value
https://t.co/5xyRQ5e350
Living among Indian villagers, something else caught his attention:
Their ability to find joy in simplicity.
No fancy gadgets. No excess. Just the essentials.
This later became Apple's philosophy:
"Simplicity is the ultimate sophistication."
A recent paper led by postdoc Kai Sun: https://t.co/vMlK7VsWC8 We improved the geographical random forest model via methods such as theory-informed hyperparameter determination and applied it to case studies in disasters and health. Code is shared at: https://t.co/32jl1zDrAj
Bayesian methods in #MachineLearning integrate prior uncertainty 🤔 with the likelihood from observed data 📊. Using techniques like Markov chain Monte Carlo (McMC), we sample the posterior distribution, accounting for burn-in periods 🔥 to reach an equilibrium state 🔄 in the chain. The resulting credible intervals 📉 provide an intuitive way to assess uncertainty in model parameters.
I’ve built a Metropolis Sampler 🧮 to demonstrate constructing a Bayesian linear regression model. It’s available in my "Applied #MachineLearning in #Python" e-book, which is free, accessible, and features well-documented workflows, downloadable code 💻, and links to my YouTube lectures 📺.
Check it out: https://t.co/8o4ykOqXMz ∀. #DataScience
R corriendo en el navegador del iPhone gracias a webr y webassembly . No hay servidor detrás ni Google colab. R compilado en webassembly. Pd: para Python está pyodide que es lo mismo
[Hilo original de @paulnovosad ]
La mayoría de los científicos ganadores del Nobel provienen de familias de élite, con padres en el percentil 87 de ingresos y 90 de educación.
Esto sugiere que el "éxito" científico está limitado por la desigualdad en el acceso a oportunidades.
🤩Scrollytelling with @quarto_pub, closeread extension and #rstats!
Great new possibilities for data-driven stories. Check out my first project: https://t.co/Esckl29fOo (on a laptop)
👀 @geokaramanis@R_Graph_Gallery
I am excited to announce the release of the first edition of my e-book, "Applied Geostatistics in Python: A Hands-on Guide with GeostatsPy".
This e-book is designed to support students & professionals learning and applying #geostatistics / #spatial #DataAnalytics through 25 chapters with well-documented, demonstration workflows that help you understand theory and apply best-practice all in modern, flexible #Python codes. By expanding from my GitHub repositories to an e-book, I aim to reach a broader audience with a dynamic and evolving educational resource.
This living document will continuously grow in response to your feedback and needs. Also, it is integrated with my online lectures, interactive dashboards, and workflows, all based on my #opensource Python package GeostatsPy and other common open-source packages. Accessible and actionable educational content!
Check it out @ https://t.co/9fvT1onXPL ∀. #DataScience
Every day, I review abstracts and papers for my PhD students, my editorial roles, and as a reviewer for various journals. I consistently notice these recurring issues in #technical#writing. Writing is a challenging yet rewarding skill, and improving it is a lifelong journey. #mentorship #ProLife #Professor
To help my students comprehend artificial neural networks (ANN), including initialization, backpropagation, and updating, I built an ANN from scratch and developed a custom interactive #Python dashboard using @matplotlib.
Now, when I teach these concepts, students engage in hands-on, experiential learning by interacting directly with the machine. It’s a fun and effective way to make complex topics more accessible!
I share it on #GitHub @ https://t.co/XDrEp91GZA ∀. #DataScience #MachineLearning
There is a hidden crisis that's slowly killing you:
Visceral fat.
I've been a health coach for +10 years.
Here are 3 ways to destroy visceral fat (bookmark this):
Last week, I added four new chapters to my recently released e-book, "Applied Geostatistics [and #Spatial#DataAnalytics] in #Python."
Having written a traditional book before, I find there's something uniquely rewarding about offering a "living document". This e-book is regularly updated and seamlessly integrated with my other educational resources, including #YouTube lectures and Python demonstration workflows on #GitHub.
It's open to all—explore it here: https://t.co/k9WQGX0jAM ∀. #DataScience