When I was first exposed to the Confusion Matrix, I was lost. And there was a HUGE mistake I was making with False Negatives. It took me 5 years to fix it. I'll teach you in 5 minutes. Let's dive in.
1. A confusion matrix is a tool often used in machine learning to visualize the performance of a classification model. It's a table that allows you to compare the model's predictions against the actual values.
2. Correct Predicitions: True Positives (TP): These are cases in which the model correctly predicts the positive class. True Negatives (TN): These are cases in which the model correctly predicts the negative class.
3. Model Errors: False Positives (FP, Type I Error): These are cases in which the model incorrectly predicts the positive class. False Negatives (FN, Type II): These are cases in which the model incorrectly predicts the negative class.
4. My Big Mistake: In machine learning we're taught to optimize for model performance. I listened. I said OK, let's optimize for F1 Score. That's the gold standard right?
5. The Problem with F1 Score: The problem with F1 is that it weights False Positives (Type 1) and False Negatives (Type 2) Errors equally. But in business this is RARELY the case. False Negatives are normally 10X to 100X more costly to a business like Netflix. Let me explain.
6. Why minimizing False Positives is worth LESS: If a predictive model (False Positive) incorrectly predicts a customer is going to leave, and Netflix decides to send them preventative actions like a discounted deal. The customer takes it and saves 10%. But they would have stayed anyway. Over a year that costs Netflix $12.
7. Why minimizing False Negatives is worth MORE: Now on the flip side, if Netflix's model incorrectly classifies some one that is on the edge of leaving as predicted to stay, Netflix does nothing. That customer leaves. Over a year that costs Netflix $120. And over the lifetime that could be $500+.
8. What they don't teach you: Expected Value (EV). I'll have a post on that soon. It's how you optimize Machine Learning models for $$$ instead of F1.
===
Ready to learn Data Science for Business?
I put together a free on-demand workshop that covers the 10 skills that helped me make the transition to Data Scientist: https://t.co/LR39RJ5XKB
And if you'd like to speed it up, I have a live workshop where I'll share how to use ChatGPT for Data Science: https://t.co/EaMpKrJiqX
If you like this post, please reshare β»οΈ it so others can get value.
#MUFC have lost the most money of the Big Six from player sales in the last 5 years with Β£192m, though this includes Β£33m on Wayne Rooney, who provided great service. Only one profit above Β£10m (Daniel James to #LUFC). Even made Β£10m loss when sold Lukaku to Inter for Β£67m.
I used work done by @recspecs730 for the Euros to estimate favourites for #AFCON2021. According to the model, the favourites for the tournament are Morocco and Senegal. The model gives each of Ghana, Nigeria and Ivory Coast a 7% chance of winning the title
@HemmenKees I think this is a result of the team set up. He is often double teamed, his full backs offer very little overlap threat to create space and the other forwards are often static. We could have Neymar/Mbappe in this team and their take on numbers would be lowπ
@nirr387@TDataScience Hi Nicklas - thanks for sharing this great tutorial. Iβve followed each line of code step by step. What is your preferred approach to handling imbalanced datasets? It wasnβt covered in th article but Iβm curious if SMOTE would have been used if needed.
@HemmenKees Kagawa, Mata, Mhikitaryn, Pogba, Sancho, VDBπ such a long list of talent that have been stifled by the rigid football of #mufc. Their teammates are not on the same wave length. Lack of coordinated movement in final third by the team makes these players look ordinary π