ex-Apple engineer gave Grok 4.6 two jobs inside Cursor, went to sleep, and opened the results live the next morning:
no babysitting, no checking every generation, just a model left running on real work for hours.
• 00:42 - reveal the website Grok redesigned overnight
• 23:21 - go from a voice prompt to a full software stack
• 45:43 - inspect the generated code + architecture
• 01:21:26 - Grok 4.6 vs Opus 5
• 01:46:31 - live PR + Cloud Agent workflow
Most coding demos test an AI for 5 minutes. This tests the thing that matters for agents: can the model keep working when you stop watching it?
SpaceXAI built Grok 4.6 specifically around longer-running agent tasks, coding and more ambitious visual work.
The endgame isn’t prompting faster. It’s giving an agent a job at night and reviewing finished work in the morning.
Worth watching before you choose which model runs your overnight agent loops.