The new 2026 FREE Data Analyst Bootcamp is live!
Here's what you'll learn:
- Data Fundamentals
- MySQL
- @Microsoft Excel
- @tableau
- Microsoft Power BI
- Python
- Pandas
- Building a Portfolio Website
- Creating a Resume
- Practicing for Technical Interviews
- @awscloud
- @Azure
- Git and GitHub
- R Programming
- @databricks
- How to use LinkedIn to Land a Job
That's a lot! All packed into one long 28 hour and 41 minute video.
Over the next year or so, I'll be creating new lessons on the following
- @alteryx
- @Snowflake
- PostgreSQL
- @duckdb
- Statistics
- and more!
Sometime in 2027 I'll release an updated Bootcamp with these included!
This Bootcamp was made with a lot of love and I hope you all learn a ton from it. The data community has given me so much so I'm glad to give back and pass it onto the next generation. Happy learning!
https://t.co/2SjI9iVxr5
INSTEAD OF WATCHING NETFLIX TONIGHT. Spend 2 hour with this. Claude AI FULL COURSE that teaches you how to BUILD and AUTOMATE anything. The people who watch this tonight will wake up tomorrow with a new skill. Watch it and Bookmark it now.
🚨 Learn Data Science for FREE (no excuses left).
Most people think you need expensive courses… you don’t.
Start with these beginner-friendly YouTube resources in 2026:
1. SQL for Data Analysis
https://t.co/10qcnjMxn9
2. Excel (Data Analysis + Dashboards)
https://t.co/iOGdRhXJNy
3. Python (Basics + Projects)
https://t.co/JW6qVUnDDa
4. Python Libraries (NumPy, Pandas, Matplotlib)
https://t.co/mhFOU6mdTw
5. Power BI (Data Visualization)
https://t.co/oX2gavWI2O
6. Tableau (Dashboard + Storytelling)
https://t.co/IBocU1Gp3g
7. Statistics for Data Science
https://t.co/b2ucYrb907
8. Data Analysis Projects (Real-world)
https://t.co/p34zcmy0eg
Stop overcomplicating it.
Pick one skill → practice daily → build projects.
Bookmark this so you don’t lose it.
RT to help someone who’s stuck.
Follow @DivyanshT91162 for more AI & Data Science content 🚀
Data Analysts!!
The next time you’re looking for a dirty dataset to clean to practice your data cleaning skill, use this prompt to generate the data from Ai.
Save for later.
“You are a data generator simulating real-world datasets for data analysis practice.
Create a dataset with the following specifications:
1. Domain / Context:
- [INSERT DOMAIN: e.g., e-commerce, healthcare, banking, education, logistics]
2. Dataset Size:
- Generate [X] rows
3. Columns (with data types and meaning):
- Provide [10–20] columns including a mix of:
- Numerical (integers, floats)
- Categorical (nominal + ordinal)
- Text fields
- Dates/timestamps
- IDs (some structured, some inconsistent)
4. Intentional Data Quality Issues (VERY IMPORTANT):
Introduce realistic “dirty data” problems such as:
- Missing values (random + patterned)
- Duplicate rows and duplicate IDs
- Inconsistent formats (e.g., dates: DD/MM/YYYY vs MM-DD-YY)
- Typographical errors in categorical values
- Mixed units (e.g., kg vs lbs, USD vs NGN)
- Outliers and extreme values
- Invalid entries (e.g., negative ages, impossible dates)
- Inconsistent capitalization and whitespace issues
- Corrupted or partially truncated text fields
- Columns with mixed data types
5. Relationships:
- Include at least 2–3 meaningful relationships between variables
- Add some noise that weakens perfect correlations
6. Output Format:
- Provide the dataset as a table (CSV format preferred)
- Include column headers
7. Additional Context:
- Briefly describe what each column represents
- Mention key data issues intentionally inserted (but do not fix them)
8. Difficulty Level:
- Make this dataset suitable for intermediate to advanced data cleaning and exploratory data analysis
Important:
- Do NOT make the dataset perfectly clean
- Prioritize realism over neatness
- Ensure the dataset looks like something collected from real operations”
The best and easiest way to generate messy dataset on Claude or ChatGPT using a well structured prompt
“Generate a messy, realistic dataset for [domain e.g. healthcare, e-commerce ]. Make it 1,000–2,000 rows with these specific problems baked in: duplicate rows, missing values, inconsistent formatting (e.g. mixed date formats, mixed casing), invalid entries (e.g. negative prices, ratings above 5), and mixed data types in the same column. The dataset should tell a story — there should be a hidden insight or business problem that can only be found through proper cleaning and analysis. Give it to me as a CSV.”
Something unique not a generic dataset you need to make research on the dataset you really want to work on as well to know what’s all about before you start.
Retweet for others to see
Try this out and thank me later