continuing to train my specialized agents with verifiable loops with Spark Loops desktop
pushed a demo script producer agent from 62% to 89%, scored by GLM 5.2, Deepseek v4, and Grok 4.5 yesterday
here is a deep dive:
Loop engineering and training your own agents with verifiable loops is about to get easy for everyone.
Here is the first look to Spark Loops desktop app, where you can:
> create your agents
> train them without any technical friction
> use them in sessions and your workflows
For the past couple days I've been turning Spark into the desktop version where you can create and train agents with verifiable loops (using multiple LLM judges) in anything you like
And then use those agents with their trained domain chips in your sessions directly from Spark Loops desktop app
Improving the work you do with them
In one example I trained an agent for writing better PRDs, distilled what separate LLMs gave for improving it, and then turned it into a PRD creator that then evolves the PRDs in the session system.
In another one I've been working on a training architect so trainings can be done better, got it to as high as it scoring 83.5 via multiple runs so that I can distill that to create better training sessions for other agents.
While trying to do all this via Telegram, the usability has been quite bad, but now everything is getting connected to a place where you can create, train, and then use the same agents in your workflows directly within Spark Loops desktop app.
6 months of work finally getting together.
A first look to Pet Boosters of Seas of Spark:
Through your pets: battles, missions, economy, and various in-game systems will be much more fun.
Pet Boosters will be earned in-game, and having a Battle Pass will get you more of them.
Join the waitlist: https://t.co/Lmn8u4zhLZ
Spark R29 is now live!
R29 consolidated accepted work across Spark CLI, Telegram Bot, Builder, Spawner, Researcher, Character, Voice, Personality, Memory, Domain Chips, and the installer surface from community PRs.
The biggest changes:
Spark CLI is safer.
R29 added stronger approval gates for risky commands, module-name sanitization, SQL allowlists, Unicode prompt-injection normalization, atomic SSH known_hosts writes, npm install-command restrictions, and better secret/path redaction.
Telegram is more reliable.
Accepted fixes improved group command parsing, API-key/stderr redaction, SQLite WAL shutdown, conversation memory races, temp-file cleanup, socket cleanup, preview URL SSRF checks, and slash-command handling.
Builder is harder to break or abuse.
R29 added provider SSRF validation, OAuth transaction safety, minimal child-process environments to avoid secret leakage, SDK module allowlists, JSON guards, and safer runtime path behavior.
Spawner got mission stability work.
R29 added lifecycle dedupe, retry tolerance, timeout settlement, reconnect jitter, scheduler in-flight dedupe, workspace containment, event cleanup, and constant-time comparison hardening.
Researcher got safer execution paths.
Accepted work covered path traversal prevention, SSRF/IP validation, stale lock recovery, subprocess timeouts, corrupt JSON guards, and public summary redaction.
Character, Voice, Personality, and Memory got durability and privacy fixes.
That includes .env exfil detection, voice SSRF protection, NaN/Inf rejection, atomic personality writes, path traversal guards, corrupt YAML warnings, log growth bounds, and safer profile loading.
The release also included smaller but important system fixes across Spark-Agent-Site, Domain Chip Labs, Domain Chip Memory, and Startup Bench.
R29’s theme: fewer leaks, fewer crashes, safer local authority, and a more reliable agent path from install → Telegram → Builder → mission → memory/persona/voice.
Spark Compete leaderboards are updated as well. Still there are more community PRs to ship in the next installers, so the current leaderboard is not final.
See you in the next updates!
Make sure to update your Spark agents 🫡
Are you a good fisherman?
Test your skills in our upcoming playtest at Seas of Spark. Don't forget to join the waitlist. The best fisherman will earn the first limited booster pets.
P.S. Invite your friends to get boosted on the leaderboards at https://t.co/Lmn8u4zPBx
While work has been ongoing with Seas of Spark, we continued to rebuild Spark Agent Harness in the past 2 weeks for
> more reliability
> curing convo hijacking bugs
> and bringing blazing fast speed compared to the older version
as well as many other updates
R28 is now live!
You can now update your Spark agents. From tomorrow onwards, we'll continue reviewing PRs.
✅ TG Raidbot activated for pirates
✅ First videos created for the vibes
Now time to load up the cannons, sharpen the swords, and generate a bunch of funny skits and videos to raid in style..
If you have cool ideas for skits and video memes to be used in raids drop them here:
Ay yo pirates, about time to use every creative tool in our arsenal to march, raid, and conquer as a community.
Join our Telegram if you have that bagworking spirit.
Channel link: Spark_coded 🦜
Hey ya, we got some news, pirates!
We gon' be reaching out to 100 Creators for Honorable Creatorships, ships designed in their imagery, to be shared with 10 of their loyal followers.
Creatorships will be NFTs that can be ridden in the Seas of Spark.
Building continues! Ahoy🏴☠️
✨updates on Spark ecosystem:
> New agent harness: the audit components that Fable 5 gave have been fixed with Codex and we are now doing the final QAs to get things live.
Harness already feels about 4-5 times faster, and much more reliable.
ETA: End of the week
> Seas of Spark: the game will be our creative way to show the potential of vibe coding, AI tools, and agents.
We will step by step integrate workflows being used in the making and marketing of the game into Spark agent.
As well as create new token utilities for Spark with it.
ETA: Aiming to start the playtests soon
> Spark Compete: PRs will start being accepted into Spark repos again right after the new agent harness release.
Thank you! 🙏
Pretty crazy that what Satya is describing here is exactly how @Spark_coded works
> domain chips: creating workflows around domain knowledge
> benchmarks: that create private evals on outputs to pave the way for improvements on the domains
> autoloops: to spawn self-improving loops on the domains while improving them on benchmarks
You can start using it here:
https://t.co/bGITaVp3Vd
To get your hands on some unique Seas of Spark ships, make sure to have some base:0x0fb07c88bc6d195c196279523957c004eb868248
And, get ready to have some fun! 🏴☠️🦜