TOP 1% AI PROJECTS that get you hired in 2026.
Projects that separate builders from learners.
1.) Terminal Agent
Build an agent that executes shell commands, reads files & debugs errors autonomously.
Target: Reach top 10 on Terminal Bench leaderboard.
Shows: You can build agents that interact with real systems safely.
2.) Open Source Agent Clone
Rebuild Hermes or Clawbot with your own improvements and benchmarks.
Target: Match or exceed original performance on eval suites.
Shows: You understand agent architecture not just API calls.
3.) Slack + AI Automation
Build a bot that triages messages, summarizes threads & triggers workflows.
Target: Deploy to 5+ workspaces, get real user feedback.
Shows: You can integrate AI into existing tools people actually use.
4.) Code Generation Pipeline
Build a system that generates, tests & refactors code with human review gates.
Target: Superset or T3.code level quality on real repositories.
Shows: You can automate development workflows not just write scripts.
5.) Generative Learning Platform
Build an AI tutor that adapts to user level, generates quizzes, tracks progress.
Target: 100+ active learners, measurable improvement in outcomes.
Shows: You can build personalized AI experiences that drive retention.
6.) Intelligent Model Router
Build a system that selects the optimal model per task based on cost, latency, quality.
Target: Reduce inference costs by 50% without quality loss.
Shows: You understand cost optimization not just model capabilities.
7.) Domain-Specific Benchmark
Create eval suites for specific use cases: legal, medical, finance or code repos.
Target: Public leaderboard, community adoption, cited by others.
Shows: You can measure what matters, not just generic accuracy.
8.) Multi-Agent Research System
Build 3+ agents that collaborate: researcher, writer, fact-checker with consensus logic.
Target: Publish one research report fully generated and verified by agents.
Shows: You can orchestrate swarms not just single agents.
9.) Production Observability Stack
Build tracing, logging, cost dashboards & alerting for deployed agents.
Target: Monitor 1000+ agent executions, catch failures before users do.
Shows: You can ship to production not just localhost.
10.) Open Source Contribution
Extend LangGraph, CrewAI or LlamaIndex with a new pattern, write docs, publish benchmarks.
Target: Merged PR, adopted by community, cited in official docs.
Shows: You are a community builder not just a consumer.
Most people stay stuck watching tutorials.
Builders get hired.
(Bookmark & Repost)
CANCEL your weekend plans.
You NEED to:
• Build a RAG system that cites sources with page numbers
• Implement hybrid search (dense + sparse) for better retrieval
• Add reranking with cross-encoders for top-10 accuracy
• Set up chunking strategies (500 tokens, 50 overlap minimum)
• Build query expansion for better recall
• Add metadata filtering for scoped retrieval
• Implement citation grounding to prevent hallucinations
• Create an eval harness with 50+ golden test cases
• Track retrieval metrics: hit rate, MRR, NDCG
• Add fallback to web search when confidence is low
• Build query rewriting for ambiguous questions
• Implement parent document retrieval for context
• Add embedding caching to reduce latency 80%
• Use colbert or late interaction for better accuracy
• Build a RAG dashboard showing retrieval quality
• Test with adversarial queries that should return nothing
• Document your chunking strategy and why it works
• Benchmark against naive RAG and show improvement
You have way too much to do.
Bookmark & Repost.
This is my personal software factory.
It turns ideas into working products while I sleep. No babysitting coding agents with prompts all day.
Cursor and Claude Code made writing code way easier. The harder problem is building a system that can context engineer and manage itself.
My factory starts with a Skill called `/factory`. It's the foreman that remembers where the project stands and sends in the right worker for the job.
Factory runs this assembly line:
1. `/factory-plan` - the interviewer
Reads the existing codebase (if there is one), extracts missing context from me via interview, and writes a product brief.
2. `/factory-plan` - the planner
The same skill turns the approved brief into small, testable features and development tasks.
3. `/factory-tests` - the professor
Before anyone writes code, every task gets an exam. This skill defines the success criteria for each task, and how the coding agent can prove to itself that what it built works or needs iteration.
4. `/factory-explain` - the presenter
Explains the plan to me like I'm 10, with visual metaphor and mermaid charts. Now that coding agents can write more code, faster than any human ever could, the new bottleneck is human understanding of the code. This skill solves that.
5. `/factory-handoff` - from CTO to SWE
This packages the brief, plan, tests, safety rails, and stop conditions into one work order. Factory uses the best models for the planning in the previous steps above, then hands the work order to a lower token usage model like Grok 4.5 for execution.
6. Cursor or Claude Code `/loop` - the coffee
The night shift picks one task, builds it, takes its exam, records what happened, iterates if needed, then moves onto the next task. If it gets stuck, circuit breakers stop it from confidently digging a deeper hole while I sleep.
7. `/factory-review` - the teacher grades the homework
The student doesn't grade it's own homework. A fresh agent that never met the builder tries to break the result. The reviewer rereads the original plan, reruns tests, and finds anything that's broken.
8. `/auto-loom-proof` - shows the evidence
Uses browser use and screen records itself performing the tests and adds an 11labs voiceover explaining what's being proven. It sends me the narrated demo video.
9. `/factory-explain` - the code
The factory updates a plain-language owner's manual explaining what actually got built. I understand my own codebase, so I can make decisions without becoming the bottleneck or outsourcing my thinking to AI.
NOTE ON BUILDING AI
Cursor and Claude Code have made writing code dramatically easier. But getting AI to work reliably and at scale for you can't be fully automated.
LLM-as-judge helps, but a judge needs a rubric, examples, and input from someone with subject matter expertise. You still need a human reviewing the work and teaching the system how to perform better.
You can check out my factory on github in the post below.
IBM just released a 1-hour course on building agentic knowledge graphs from scratch:
• 00:00 - Introduction to knowledge graphs
• 05:35 - Building your first agentic graph
• 19:59 - Agentic memory powered by graphs
• 30:39 - Graphs for multi-agent orchestration
This 1-hour watch will replace 10 paid courses on agentic engineering.
Watch it today, then learn how to become a knowledge graph engineer in the article below.
I created a handbook to help you learn AI agents.
It gives you:
• Must-know AI GitHub repos.
• Free courses to master AI agents.
• Papers to understand AI fundamentals.
• Curated videos to learn AI agent foundations.
• Books to get started with AI agent engineering.
• Condensed guides to broaden your AI agent knowledge.
(24 HOURS ONLY!!!)
To get it for free:
1 Follow @systemdesignone [MUST]
2 Like & Retweet to get DM
3 Reply "Handbook"
Then I'll DM you the details.
After more than 1 year of writing all day every day, my books are finished.
Everything I learned from making my first $1M from my personal projects.
13 free short books, based on my daily blogs since 2018.
My sabbatical is over. Cannot wait to build new businesses again.
Keeping up with AI news is becoming a full-time job.
So my friend @ivan_bezdomny built HuggingNews, an AI-curated feed that surfaces the news actually worth reading. Soon, it will even personalize the feed using your Hugging Face profile. Been using it for weeks!
Bookmark it, or ask your agent to send you the top 10 stories every morning or night. Less noise. More signal. More building!
https://t.co/h0xpkq1aHO