Sonnet 4.5 + MCP + n8n = AI Content Infrastructure that replaces $12K+/month ghostwriters...
The 3-layer system that cloned my writing voice and generated 25M organic views
→ No more 2-3 week turnarounds for 5 basic posts
→ No more robotic AI that screams "ChatGPT wrote this"
→ No more inconsistent posting killing 6+ months of momentum
→ No more ghostwriters charging $4K+ and still missing your voice
Just upload 20 posts → autonomous 3-layer content engine running 24/7.
Here's how it works:
→ Voice Intelligence Layer (learns your exact patterns from 20+ posts)
→ Psychological Conversion Engine (ICP profiling + 40-40-20 framework)
→ MCP Automation Infrastructure (YouTube sentiment + viral trigger analysis)
→ Multi-Platform Generator (LinkedIn, Twitter, newsletters, YouTube)
→ Extended Thinking Protocol (ensures 100% voice consistency)
→ 24/7 Systematic Deployment (zero manual content creation)
Built with Sonnet 4.5 extended thinking.
Runs 24/7 without supervision.
10-minute setup per layer.
The system behind 25M organic views and 80K followers.
Want the complete infrastructure?
Like + comment "CLONE" + repost, and I'll DM it to you.
(must be following)
WE'RE HIRING FOUNDING ENGINEERS AT WTF.
We’re building one of the most exciting projects yet at WTF, something that will shape us as a brand, as a platform, and (hopefully) create a far wider impact with and for the youth of India.
The core team will work directly with me and a handful of the best folks from WTF. This is on-site in Mumbai, full-time. Ground floor. High stakes. Real impact.
Here’s who we’re looking for:
• 4–5 Founding Engineers ready to roll up their sleeves.
• Builders who’ve taken products from 0 → 1 and scaled them.
• Strong fundamentals, obsessed with end-user experience.
• Comfortable making smart trade-offs across scale, cost, and latency.
• AI-native thinkers who can straddle frontend, backend, ML, infra, and stitch it all together.
• The curious ones who tinker, break, hack, and rebuild. If you’ve been exploring with LLMs, Cursor, or AI-first tools, even better.
Show us your GitHub repos, the side projects, experiments, and crazy hacks that reflect your curiosity.
If the idea of building something transformative from day zero excites you...
Email [email protected] or drop your GitHub link and profile.
We’ll reach out if it’s a potential fit.
Let’s go!
— Nikhil
New post re: Devin (the AI SWE). We couldn't find many reviews of people using it for real tasks, so we went MKBHD mode and put Devin through its paces.
We documented our findings here. Would love to know if others have had a different experience.
https://t.co/DDqzoAXKkl
I will teach Large Language Model Systems again in Spring 2025. (11868 for CMU folks) The course syllabus (tentative) is online at https://t.co/6AmKDbdHt6
CMU ppl are welcome to enroll. For others, I will release the materials online. Or apply to CMU LLM/GenAI certificate program
NVIDIA's $7B Mellanox acquisition was actually one of tech's most strategic deals ever.
The untold story of the most important company in AI that most people haven't heard of
1/12
LoRA Learns Less and Forgets Less: When I saw a new, comprehensive empirical study of Low-Rank Adaptation for finetuning LLMs, I had to read it! Here are the main takeaways.
This study aimed to compare LoRA to full finetuning on two different target domains: programming and mathematics (rather than the usual general instruction following tasks.) Moreover, the authors also compared instruction finetuning and continued pretraining scenarios.
LoRA vs full finetuning? It's maybe as expected: It all comes down to a learning-forgetting trade-off. Full finetuning results in stronger performance on the new target domain, whereas LoRA maintains better performance on the original source domain. Intuitively, I suspect this is simply a side effect from LoRA changing fewer parameters in the model -- the goal of LoRA (as its name implies) is a low-rank adaptation, that is, not substantially modifying all model parameters.
Nonetheless, it's really nice to see this all laid out and analyzed in great experimental detail. (The experiments were done with 7B and 13B Llama 2 models).
In practice, it's also often not a question whether to use full finetuning or LoRA as the latter may be the only feasible one due to its memory savings and lower storage footprint.
RELEASE DAY
After almost 10 years of hard work, tireless research, and a dive deep into the kernels of computer science, I finally realized a dream: running a high-level language on GPUs. And I'm giving it to the world!
Bend compiles modern programming features, including:
- Lambdas with full closure support
- Unrestricted recursion and loops
- Fast object allocations of all kinds
- Folds, ADTs, continuations and much more
To HVM2, a new runtime capable of spreading that workload across 1000's of cores, in a thread-safe, low-overhead fashion. As a result, we finally have a true high-level language that runs natively on GPUs!
Here's a quick demo:
Our computer vision textbook is released!
Foundations of Computer Vision
with Antonio Torralba and Bill Freeman
https://t.co/We0ZSJzkle
It’s been in the works for >10 years. Covers everything from linear filters and camera optics to diffusion models and radiance fields.
1/4
Foundation models are well-established in vision and language, but time series forecasting has lagged behind - it still relies on dataset-specific models.
Meet Lag-Llama: the first open-source foundation model for time series forecasting!
CUDA-MODE Lecture 3: Getting Started with CUDA
Video: https://t.co/ZH0cET5QZU
Notebook: https://t.co/auD6GA9gjr
🏎️Cuda intro for everyone with a Python background! @jeremyphoward builds the kernels 1:1 in python first (with blockIdx & threadIdx) ->then converts them to cuda C.
CUDA-MODE 3: Getting Started With CUDA
How do you actually write a kernel and call it from Python? How do you test and debug your code?
Speaker: @jeremyphoward
Sat, Jan 27
12:00 PM PST (Bay Area) / 9:00 PM CET (Berlin)
Live on discord: https://t.co/A2fbds6U5q