@formalaai@matviy As a heavy cursor and ChatGPT user, I’m not switching anytime soon even with this news. Cursor is an incredible product once you know how to wield it.
Why hosting your SaaS in a VPS is a must in 2026
Whether you’re building one or multiple SaaS products, having a VPS is a must
here’s why:
> Save costs on tools
Probably the biggest advantage a VPS has over normal hosting is the ridiculous amount of money you’ll save
Normal hosting (vercel, railway, render, d!bs, etc) charge you per usage:
Frontend & backend hosting
Db hosting
Redis and workers are charged separately in most cases
The more your app grows the bigger the costs u incur
The more apps you have the higher the charges
VPS hosting:
You rent a server/ computer and it becomes your computer in the cloud
Costs start at $5-7/month to several hundred depending on the specs of your ideal remote computer
I’m renting a 2 vcpus, 4gb ram, 80gb ssd on Hetzner for $26/month
This means you can host your website, db (mongodb, Postgres, nosql, etc), redis, workers, multiple SaaS, and any other services you need at no additional cost.
> Freedom
A VPS gives you total control which i think is its selling point in the age of AI agents.
For example, besides hosting my SaaS and mongodb entirely on the vps, I’ve also installed Hermes in there.
It acts as my apps supervisor. It reads my db to uncover angles, data correlations, assess leads, create reports and send to telegram, and so much more
Yes, I’m fully aware you can do this on normal hosting but separately.
But the advantage here is it’s all happening in the same box that runs 24/7 uninterrupted
So many possibilities!
The only downside of a VPS is you have to handle everything yourself: security, deployment, servers, logging, monitoring dashboards, etc
Luckily, your favorite agent can help you set this up
I’m also down to help where I can, just shoot me a dm.
> be Long Term Capital Management (LTCM) in ‘98
> a hedge fund built by former Alpha traders from Salomon bros.
> like how Anthropic was born out of OpenAI
> talk shit, do what u want cuz you’re kings of Wall Street
> have Nobel Prize winners Robert Merton & Myron Scholes on the team
> make +43% in 1995 and +41% in 1996
> convince Wall Street your models have cracked markets
> start leveraging the fuck out of everything
> $4.8B in capital
> ~$125B in assets
> ~$1.3 TRILLION in derivative notional exposure
> roughly $30 borrowed for every $1 of capital.
> MADNESS.
> strategy is basically “these tiny price differences will eventually converge”
> then Russia defaults in August 1998
> markets go completely fucking insane
> the relationships your models considered almost impossible to break… break
> how? the best math in the world can’t factor in human stupidity
> LTCM loses ~44% IN ONE MONTH
> capital collapses from ~$4.7B to under $1B
> now they can’t even dump their positions without potentially blowing up the market
> it’s now September 1998
> the NY Fed gets 14 banks in a room
> banks inject $3.6B
> take control of LTCM
> not because LTCM was some random hedge fund
> but because everyone realized liquidating $125B of positions at once could fuck everyone else too
> the models weren’t “wrong” because they couldn’t do math
> they were wrong because the assumptions about how markets behave completely broke
> moves were described as 10-sigma events
> events their historical models considered so absurdly unlikely that they basically lived in the “this should never happen” bucket
> then it happened
> finance markets blow up worldwide
> banks remain holding a fossil
He was deemed the oracle of AI, with visions of buying galaxies and colonizing planets. The forces that unraveled his fund were ones he didn’t see coming. https://t.co/TwjgnNp0Os
Hermes is bleeding tokens by default!
Hermes requests including startup (approx. 16k-30k tokens) hits the model’s API with ALL plugins, tools, and context-compress on long sessions by default.
This is EATING up tokens that would be otherwise used to solve problems.
Here’s how to solve this:
Tell Hermes to route auxiliary tasks to the keyless free provider opencode-free (ships with Hermes), instead of hitting your paid model.
Trim unused toolsets to the ones u use per profile. The system prompt ships the schema of every enabled tool on EVERY API call. You can always add more when the need arises.
He was deemed the oracle of AI, with visions of buying galaxies and colonizing planets. The forces that unraveled his fund were ones he didn’t see coming. https://t.co/TwjgnNp0Os
I lost a friend a month ago abruptly.
Nice guy, came from money, had his own business, family has businesses
But..
He was super conservative with money.
I remember on his birthday in 2023, we were out and he refused to treat himself that I ended up spoiling him a bit to celebrate his birthday.
He passed away never having travelled abroad, never owning the fancy car despite the family owning a dealership, etc
It was a real eye opener! We live like we own tomorrow so we postpone living, yet it’s the only reason we are here really
I sold my house to buy a Ducati.
Yes. That Dukaan founder/CTO sold his house for a f*cking Ducati.
Let me tell you about yesterday first.
An old man walked up to me at a signal. Turned out he’d followed MotoGP his whole life. Rattled off riders, their bikes, championships, decades of it. Then he went quiet and asked if he could just twist the throttle once.
Once.
I said yes. The look on his face, I’ll never forget it.
Later, a guy in his thirties stopped me. A total bike nerd who knew every spec of the XDiavel better than I did.
Maybe it was his dream bike. He asked for a ride, and I gave him one. I’ve never seen someone that happy.
Now here’s how I got here.
I Bought a house in Mumbai in 2018 for ₹1.8 Cr. Later when we Started Dukaan in 2020, i moved to Bangalore.
The house stayed in Mumbai. I didn’t.
But the EMI followed me everywhere.
Every month, I told myself the lie millions of us tell ourselves: “It’s an investment. It’ll appreciate. Hold long enough and it’ll make sense.”
i kept paying over 1 lakhs every month to the bank. mostly interest.
For a house I barely lived in.
An asset 1,000 km away. A present I kept postponing for a tomorrow that never came.
At some point, “investing for the future” becomes a very expensive way of refusing to live in the present.
Meanwhile, The endless grind of Building a startup, the middle-of-the-night anxiety, the weight of keeping the machinery moving was keeping my head spinning at all times.
A 20-year mortgage on a piece of concrete wasn’t going to fix my head.
I needed a release valve.
So I sold the house and bought a Ducati XDiavel.
And let me tell you, it’s a good f*cking bike. 1,262cc V-twin. 156 hp. Brembo brakes.
The second the engine fires.
The sound.
The first twist of the throttle.
When it roars to life, the happiness hits so hard I cry inside my helmet.
A quick ride around ORR at 1 AM the noise inside the head finally goes quiet.
When I think about that old man.
Maybe I’ll have the money at 60 too. But will I still be able to feel it? Or will I just know the specs, and ask a stranger for one last twist of the throttle?
Buying the bike isn’t the stupid decision.
The stupid decision is not buying it, and realizing too late that the moment already passed.
I spent years building a future. Maybe it’s time I start living in it.
And for you; Stop going into debt for your future. Go buy your present. Get into debt if you have to..
Get the bike. Get the ridiculous music system. Build the home lab. Book the trip.
Buy the thing that makes you cry inside
Hermes is bleeding tokens by default!
Hermes requests including startup (approx. 16k-30k tokens) hits the model’s API with ALL plugins, tools, and context-compress on long sessions by default.
This is EATING up tokens that would be otherwise used to solve problems.
Here’s how to solve this:
Tell Hermes to route auxiliary tasks to the keyless free provider opencode-free (ships with Hermes), instead of hitting your paid model.
Trim unused toolsets to the ones u use per profile. The system prompt ships the schema of every enabled tool on EVERY API call. You can always add more when the need arises.
Hermes command cheat sheet. Save this. You’ll probably need it again.
The useful commands are spread across Desktop, CLI, messaging, and the terminal, which makes it easy to mix up what works where.
So I put the ones worth knowing into one reference, organized by session control, active work, models, skills, automation, recovery, and more.
Bookmark it and keep it around.
Hermes had another ridiculous week.
Last week was huge. Somehow, it did not slow down.
In the last seven days, Hermes learned to control the browser you are actually looking at, Bot Mode turned into a multi-machine agent fleet, conversations gained live interactive UI, Cron got persistent memory, and `/review` gave finished work an independent second set of eyes.
And those are just the headliners.
Here are the biggest Hermes changes from August 17-23:
▸ Hermes can now use the browser you are actually looking at:
The in-app Preview browser is no longer something Hermes can only open and read.
Hermes can inspect the page, click, type, scroll, press keys, navigate, reload, and even annotate an element so you can see exactly what it is targeting.
That means Hermes can work inside the visible browser session you are already signed into while you watch it happen.
This opens up a ridiculous number of workflows.
▸ Bot Mode became a real multi-machine agent fleet:
A lot of what was coming together last week became real infrastructure this week.
Desktop can now give you one Bot Mode roster across registered Hermes gateways while keeping every bot tied to the machine that actually owns it.
Local. Remote. SSH. Hermes Cloud. Different machines.
Remote bots can open their real Bot Chat without moving your normal Sessions workspace.
Bots also gained `message_agent`, giving them an actual agent-native way to message teammates instead of constructing shell commands.
And bots across different Desktop connections can message each other too.
Group rooms also gained real threads plus images, PDFs, files, paste, and drag-and-drop attachments.
Bot Mode is starting to look a lot less like “multiple profiles” and a lot more like an actual team.
▸ Your Hermes conversation can now become an interface:
Plugins can render live UI directly inside the conversation.
Not a screenshot.
Not a link to another app.
Actual interactive UI inside the transcript.
And it does not have to be one-way.
Interacting with that UI can send a real turn back to Hermes behind the scenes, let the agent do more work, and update what you see.
Chat does not have to end in text anymore.
▸ `/review` gives your agent an independent second set of eyes:
Your main agent finishes the work.
Run `/review` and Hermes launches a separate background reviewer with its own tools to inspect the actual work behind the conversation.
Code. PRs. Docs. Research. Other referenced artifacts.
The reviewer can work under the repo’s own project instructions and can be told which skills the primary agent was using.
You can also give Review its own model.
One model does the work.
Another independently checks it.
The findings return to the original session so the main agent can respond, defend the work, or fix what it missed.
This is a big one.
▸ Cron became much more than “run this prompt at 8 AM”:
Cron agents can now use Hermes persistent memory, including MEMORY.md and USER.md.
Individual jobs can also pin their own reasoning effort, so a lightweight recurring check does not need the same reasoning budget as a scheduled deep analysis.
And scheduled output can now be delivered directly into a bot’s canonical Bot Chat as a real incoming turn.
The bot reads it and responds.
Now start combining scheduling, memory, Bot Mode, and unattended work.
That gets interesting very quickly.
▸ `hermes update` now checks whether the update actually worked:
This one is less flashy, but it matters.
Hermes has been rebuilding the updater around the reality that one installation may have multiple profiles, gateways, services, and runtimes running at once.
`hermes update --plan` can inventory that fleet before touching anything.
Updates generate structured receipts.
Running gateways can report which code they are actually serving.
And the restart phase now has to account for the runtimes that were in the plan.
If Hermes expected something to be updated and it was missed, stale, or down, the updater can fail visibly instead of quietly declaring victory.
That is the kind of boring infrastructure improvement you really appreciate the first time something goes wrong.
▸ Long-running jobs got a serious reliability pass:
The old default 500-turn ceiling is gone.
Hermes now runs with unlimited turns by default unless you choose to set a cap.
New stall guards also look for agents repeatedly making the same tool call or announcing that they are about to continue and then stopping.
And if Hermes runs the same tool again and gets the exact same giant result, it still executes the tool fresh, but it can feed the model a tiny reference to the previous result instead of dumping another 20K-50K characters into context.
MCP results also got tighter context handling.
The goal is pretty simple:
Let long jobs keep working without letting the agent quietly spin its wheels or bury itself in repeated context.
▸ Fresh Hermes installs need fewer API keys before they can do useful work:
Hermes now has keyless paths for web tooling across multiple supported providers.
There is also a new `opencode-free` provider for OpenCode’s free tier that does not require an API key or OpenCode account.
It appears in the normal model pickers too.
Install Hermes.
Pick an available free model.
Give it a job.
The distance between “I just installed Hermes” and “my agent is doing useful work” keeps getting shorter.
▸ Repositories can now bring their own Hermes skills:
A repo can include project-specific skills under `.hermes/skills/` or `.agents/skills/`.
Inside that project, those skills can take priority over lower-level versions of the same skill.
Hermes also puts a trust boundary around this.
The repo has to be trusted, project skills are scanned when loaded, and an unknown repo does not silently get permission to inject its own agent instructions.
For teams building repeatable Hermes workflows around a codebase, this could become extremely useful.
Also worth knowing:
→ Desktop voice can use supported STT/TTS providers directly through the active profile.
→ Eligible Codex GPT sessions now default to 272K context, with explicit `-900k` variants when you actually want the larger window.
→ `/model` got fuzzy search.
→ Desktop got a command palette.
→ `/status` got much more useful session information.
→ `clarify` can ask several independent questions in one batch instead of interrupting you one question at a time.
→ Hermes can apply Desktop layout presets itself.
→ Plugins can be installed from Git or `hermes://` links with a review step before anything gets installed.
Last week’s story was Hermes learning to keep working, delegate work, supervise agents, and coordinate across machines.
This week is what happens when those pieces start connecting.
The agent can use the browser you see.
Bots can work and communicate across machines.
Scheduled jobs can remember, wake bots, and hand work back into the team.
Agents can build interfaces inside the conversation.
And when the work is finished, another agent can independently inspect it.
Hermes is starting to look less like one very capable agent with a long feature list and more like an operating layer for a team of agents that can actually keep working together.