AWS's September 9 announcement of a 90-minute Lambda timeout didn't land as a clean win. The reaction was more mixed than I expected.
One AWS engineer called it the end of the duct tape people used to work around the old limit. Others pushed back, if your function genuinely needs 90 minutes, ECS Tasks or AWS Batch were probably the better home all along.
I find that debate more interesting than the release note itself, whether you patch the serverless model to hold longer jobs, or admit the job was never serverless shaped to begin with.
See for yourself here https://t.co/l2n1p9yInA
The engineering deep dive on Netlify Edge Functions running on Unikraft is live. Find out (in detail) how we get a microVM booted with customer code on disk in 2ms at P99.
Full post: https://t.co/b77BFlyjeZ
After announcing the Netlify collab yesterday, a lot of people were interested in how they can put compute in the request path for edge functions. Good thing I wrote a case study in close collaboration with their platform engineers! Linked below, full of juicy technical detail :)
Link to the full story here: https://t.co/KhpqdnPlBm
BIG NEWS: @Netlify's entire Edge Functions product now runs on Unikraft. Every tier, all their edge compute, more than 20 BILLION requests a month for them.
Full story on how and why:
Link to the full story: https://t.co/TaMilv8dsl
Release 13 is out! It contains RM64 hosts, kernel-less images, Karpenter provider, JavaScript SDK, flexible networking, new dashboard, etc. I really like the short and sweet videos Felipe has been doing lately to walk everyone through the release 👇
Today we are releasing version 13 of our platform! It is absolutely PACKED with SOTA features:
ARM64 hosts, kernel-less images, a Karpenter provider, a JavaScript SDK, way more flexible networking, and a much more useful dashboard. Felipe called it the biggest release since we shipped GPUs, and I'm not about to argue with him.
I’m going through all of them in the next few days in more detail 🙌
Full release notes here https://t.co/i3wXDOt6qH
If memory is scarce and expensive right now, packing more onto each server matters more than it used to. On our platform that's 100K+ VMs on a single box, hardware-isolated, not just densely containerized. Idle instances hold no memory at all, so you're not paying the RAM tax on capacity nobody's using.
RAM prices are killing everyone… Amazon raising prices for their consumer range, Samsung expecting prices to keep rising well into 2027, etc. Some of this might be excuses for margin improvements but a large part is definitely real and demand driven.
If RAM stays this scarce, memory sitting in an idle VM becomes increasingly valubable. On our platform an idle instance holds **zero** memory, not just "a bit less". We suspend it to a snapshot and bring it back in under 10ms.
Source: https://t.co/UKiqQOmwJN
"Being able to spin up a completely isolated instance with your code frozen mid-execution, in a couple of milliseconds is a very powerful primitive."
That's not us saying it, that's Indices' co-founder Lorian Richmond, about what our platform does for them.
Of course, this doesn't stop at web-request APIs. Same primitive, same milliseconds, works for spinning up agent sandboxes, racing a few headless browsers to see which one loads first, or forking a database so someone can poke at it without touching production.
Full story at https://t.co/D2unzINCo9
Every one of Indices' connector calls runs in its own microVM, with its own hypervisor boundary, no shared kernel with anything else on the box. That matters more than usual here because those calls are handling other people's credentials, MFA codes, TOTP, login sessions, etc.
Their co-founder Lorian Richmond put it simply: "good security goes hand-in-hand with simplicity."
Full story at https://t.co/D2unzINCo9
Keeping a warm pool running 24/7 so cold starts don’t ruin your API is a pretty expensive way to guarantee speed.
That was Indices' setup before Unikraft: Instances billed around the clock whether they were doing anything or not, plus the orchestration logic to keep the pool healthy. With 6ms boots, capacity just follows demand. Nothing pre-provisioned & nothing idle.
Full story at https://t.co/D2unzINCo9
Unikraft’s snapshotting feature cut Indices' per-call latency by about 50%, just by skipping unnecessary work that used to happen on every single request.
Their connectors do a lot before they even get to the actual web request: They resolve code, set up proxies, initialie clients, etc. On Unikraft we freeze a fully initialised instance mid-execution and resume from exactly that point, so all of that setup happens once, not on every call.
Same idea behind our other scale-to-zero work, just pointed at their specific bottleneck :)
Full story at https://t.co/D2unzINCo9
2,000 milliseconds down to 6! That’s how much Indices (https://t.co/WfdpjS1rnG) were able to reduce their cold starts by switching to Unikraft. They build APIs that let you automate web workflows, real-time enough for voice agents, so every millisecond in that call counts.
Before us, a 2 second cold start was a non-starter for something sitting in a live API call. So they kept a warm pool running around the clock to deal with that. Now they cold boot straight inside the request, like magic.
Full story at https://t.co/D2unzINCo9
Shoutout to Lorian (https://t.co/3fagBuxvyk) and Mandeep (https://t.co/DR0OXSptU3) for being such a pleasure to work with
The web is super messy, and agents need to interact with real web pages all the time. Not everything is a nice API, MCP, or CLI unfortunately. Normally, booting browser instances takes 10-30 seconds, which is extremely slow for agentic workflows at scale ☹️ Additionally, they take up GBs of memory, making them expensive to run.
On Unikraft they come up in under 10ms, and scale back down to zero the the same time frame when they’re done being used, and consume no compute or memory when idle. Want to see how this works? We have a < 3min video explaining everything: https://t.co/tHvHw5ZxBx
Something that I forgot to mention about running DBs on Unikraft (and workloads in general, I suppose) is that you can basically run Unikraft anywhere. The compute can run on standard AWS EC2 instances, including both regular virtualized and bare-metal instance types, but the same setup also works on GCP and Azure equivalents, so it isn't locked into a single cloud provider.
You can even bring you own metal or virtual machines and point them directly at a Unikraft cluster. See it in action! https://t.co/hzOSxiFMUi
I know I have been posting about our k8s implementation quite a bit this week and peopl have asked me “cool stuff, but does it also work with multiple nodes?”
There is now a section in the video where we show this: Our engineer replicates 10 pods, and all replicas come up with no observable delay. So no, we are not constrained to just single pod wake-ups! See for yourself here https://t.co/LIdIgS0FDX