@mitchellh How are you guys handling PTY resizing when multiple clients are connected? For example, does the PTY resize based on the dimensions of the web view or desktop app? And what happens when multiple clients with different viewport sizes are attached at the same time?
How do runtimes like Go and Tokio run huge numbers of concurrent tasks without creating one OS thread per task?
They use a scheduler and a relatively small pool of worker threads, commonly around one worker per CPU core.
When you call go foo() or tokio::spawn(foo()), foo becomes runnable and is usually placed in a worker’s queue.
The current goroutine/task normally continues, but the exact execution order is not guaranteed. Each worker prefers its local queue because it’s fast.
If a worker runs out of work, it can steal tasks from another worker’s queue. This balances the workload and keeps CPU cores busy.
When a task waits for I/O, the worker can run something else. Blocking work needs special handling: Go may use additional OS threads, while Tokio provides a separate blocking task pool.
The details differ, but the core idea is scheduling + queues + work stealing.
“Stack is faster than heap”, but both are just RAM.
So what actually makes the stack faster?
1. Allocation.
Stack: sub rsp, 32, essentially a pointer bump.
Heap: malloc() may search free lists, synchronize, and sometimes interact with the kernel.
2. Addressing.
Stack locals are often accessed as [rsp+offset] or [rbp-offset]. The base address is already in a register.
Heap objects often require loading a pointer first, then dereferencing it.
3. Cache locality.
The stack is typically accessed in a small, contiguous region, so recently used stack data tends to stay hot in the CPU cache.
Heap allocations can be scattered across memory.
It’s not “stack memory is faster than heap memory.”
It’s allocation cost + locality + access patterns.
A contiguous heap array walked linearly can be just as fast as stack memory.
A linked list can be slow regardless of where it lives.
A short history of how we kept finding new ways to make computers faster, from pipelining and out-of-order execution to SIMD, multicore, SMT, GPUs, and specialized accelerators. https://t.co/ckDSSAs36a
“Stack is faster than heap”, but both are just RAM.
So what actually makes the stack faster?
1. Allocation.
Stack: sub rsp, 32, essentially a pointer bump.
Heap: malloc() may search free lists, synchronize, and sometimes interact with the kernel.
2. Addressing.
Stack locals are often accessed as [rsp+offset] or [rbp-offset]. The base address is already in a register.
Heap objects often require loading a pointer first, then dereferencing it.
3. Cache locality.
The stack is typically accessed in a small, contiguous region, so recently used stack data tends to stay hot in the CPU cache.
Heap allocations can be scattered across memory.
It’s not “stack memory is faster than heap memory.”
It’s allocation cost + locality + access patterns.
A contiguous heap array walked linearly can be just as fast as stack memory.
A linked list can be slow regardless of where it lives.
@adityazero_@SakshiSugandhi Do you think people are raising PRs without fully understanding the system, or is this simply an increase in productivity for people who already know what they’re doing?