Once I have reverced the question and phrased it like "if I generate a bunch of random variables with some volatility-will the mean absolute return of this distribution be x?" , then he finally got the right idea.
What is ๐๐ผ๐ป๐๐ถ๐ป๐๐ผ๐๐ ๐ง๐ฟ๐ฎ๐ถ๐ป๐ถ๐ป๐ด (๐๐ง) in MLOps and what steps are needed to achieve it?
CT is the process of automated ML Model retraining in Production Environments on a specific trigger. Letโs look into some prerequisites for this:
1๏ธโฃ Automation of ML Pipelines.
๐ Pipelines are orchestrated.
๐ Each pipeline step is developed independently and is able to run on different technology stacks.
๐ Pipelines are treated as a code artifact.
โ You deploy Pipelines instead of Model Artifacts allowing Continuous Training In production.
โ Reuse of components allows for rapid experimentation.
2๏ธโฃ Introduction of strict Data and Model Validation steps in the ML Pipeline.
๐ Data is validated before training the Model. If inconsistencies are found - Pipeline is aborted.
๐ Model is validated after training. Only after it passes the validation is it handed over for deployment.
โ Short circuits of the Pipeline allow for safe CT in production.
3๏ธโฃ Introduction of ML Metadata Store.
๐ Any Metadata related to ML artifact creation is tracked here.
๐ We also track performance of the ML Model.
โ Experiments become reproducible and comparable between each other.
โ Model Registry could and in some cases should be treated as part of ML Metadata Store.
4๏ธโฃ Different Pipeline triggers in production.
๐ Ad-hoc.
๐ Cron.
๐ Reactive to Metrics produced in Model Monitoring System.
๐ Arrival of New Data.
โ This is where the Continuous Training is actually triggered.
5๏ธโฃ Introduction of Feature Store (Optional).
๐ Avoid work duplication when defining features.
๐ Reduce risk of Training/Serving Skew.
๐ ๐ ๐๐ต๐ผ๐๐ด๐ต๐๐ ๐ผ๐ป ๐๐ง:
โก๏ธ Introduction of CT is not straightforward and you should approach it iteratively. The following could be good Quarterly Goals to set:
๐ Experiment Tracking is extremely important at any level of ML Maturity and the least invasive in the process of ML Model training - I would start with ML Metadata Store introduction.
๐ Orchestration of ML Pipelines is always a good idea, there are many tools supporting this (Airflow, Kubeflow, VertexAI etc.). If you are not doing it yet - grab this next, also make the validation steps part of this goal.
๐ The need for Feature Store will wary on the types of Models you are deploying. I would only suggest prioritizing it if you have Models that perform Online predictions as it will help with avoiding Training/Serving Skew.
๐ Donโt rush with Automated retraining. Ad-hoc and on-schedule will bring you a long way.
Let me know your thoughts! ๐
--------
Follow me to upskill in #MLOps, #MachineLearning, #DataEngineering, #DataScience and overall #Data space.
Also hit ๐to stay notified about new content.
๐๐ผ๐ปโ๐ ๐ณ๐ผ๐ฟ๐ด๐ฒ๐ ๐๐ผ ๐น๐ถ๐ธ๐ฒ ๐, ๐๐ต๐ฎ๐ฟ๐ฒ ๐ฎ๐ป๐ฑ ๐ฐ๐ผ๐บ๐บ๐ฒ๐ป๐!
Join a growing community of Data Professionals by subscribing to my ๐ก๐ฒ๐๐๐น๐ฒ๐๐๐ฒ๐ฟ.
On this point, I think Euan has a lot to share + it is so expensive that I expect only practitioners can afford it.
I'd say main points from him:
- bro just use statistics
- monte carlo monte carlo monte fucking carlo (I agree)
- understand your edge and trade accordingly
The most legendary investor of all time: Warren Buffett
In 1987, Warren wrote a letter to Berkshire shareholders covering a variety of vital long-term investing topics.
In just 16 paragraphs he put together a masterclass in business & investing.
Here's a breakdown of each one:
Finished writing a full guide on pairs trading, specifically finding great portfolios. Itโs a pretty comprehensive dive into the topic.
Pairs trading is an area I worked on for a while, but long enough ago that Iโm happy to share my insights on it
https://t.co/86iS9JQqgA
Machine learning is famous for its open resources.
We collected 3 free courses about ML:
โช๏ธ Introduction to Machine Learning, @MIT
โช๏ธ Mathematics for Computer Science, @MIT
โช๏ธ Practical Deep Learning, @fastdotai
๐งต
2 PDFโs of Charlie Munger that every Investor shouldโve read:
- Charlie Mungerโs Harvard Speech โThe Psychology of Human Misjudgmentโ
- โThe Art of Stock Pickingโ
You can find both PDFโs for free on my Website:
https://t.co/TtH93zeco0
Very nice illustration of the Data Pipeline by Semantix. It may provide some insights into understanding data pipelines.
Subscribe to our system design newsletter to get a Free System Design PDF (158 pages): https://t.co/st1UJgX3QR
The article provides a comprehensive guide for using the wavelet transform in machine learning with examples and code snippets.
https://t.co/PJoRou3LpB
2) @RunWayML:
AI technology allows you to create your own video or movie just by writing an idea. Bring your ideas to life.
With RunWayML, you don't need any filmmaking or editing experience - just your imagination.
AI technology turns your vision into a video.
Matrix multiplication is not easy to understand.
Even looking at the definition used to make me sweat, let alone trying to comprehend the pattern. Yet, there is a stunningly simple explanation behind it.
Let's pull back the curtain!