@immanuel_vibe >kernel never spends a cycle
Except the driver runs in kernel...
eBPF is great, but a lot of these solutions are just workarounds for having a monolithic kernel vs micro kernel. We run most of these same features directly from userspace with much lower overhead than eBPF.
@iximiuz Very nice explanation of the startup internals. Another fun fact, main is just a linker symbol, at least with 10 year old compilers it doesn't even need to be in a code segment. https://t.co/pMNo88oTUC
@techneo@TrisH0x2A Its actually slow. The multiply alone is the same or more cycle cost as tzcnt, then there is a memory lookup which can be ~500x slower if not in cache.
Unless you mean the complier replacing it with 1 instruction, then yes that's awesome.
@CalebChamberla6 Had my first order ship from there a week ago or so. Somehow the cut quality is better than the last time I got the part. Amazing job having the cuts dialed from day 1.
@BenjDicken Off by ~5x in some cases. That is likely lightly loaded latency on nvme. Under load reads that hit the nand on it are way worse than 60us, and your kernel + pci transaction time on gen3 is close to 60us. Gen5 this is close. Also EBS is a lot faster than that :)
@jhleath >ebs drives are like the disks that are attached to your laptop. they don't have any logic inside of them to allow multiple machines to coordinate on the same set of data.
https://t.co/oY8RwUZgts
@Sirupsen The 1-2us "SSD" speeds look like block cache times. I'm pretty darn sure you're not getting through PCI + interrupt stack in 2 us. Hell, if your machine has deep sleep power modes on, the interrupt wake latency is over 100us.
@max0x7ba@ChShersh I feel like the talk probably about the lock free queue itself where you already have data in L1d and want to shove a pointer through a queue. The cache coherency there is usually high if there is contention with traditional CAS based approaches.
I want to find the talk!
@ChShersh@thedcoffman What CPU is that running on? Sub 10ns implies Shared L2 unless it's some 5ghz monster? And maybe doing it without locked instructions for the write? So no CAS and opportunistic reads?
@Adriksh "no memcpy" -> writes a bytewise memory copy manually and stores the whole data in the ring.
The way to do this right is to pair it with an object pool and pass the pointers through the ring. That is zero copy. Of course your object pool is likely to use the same structure.
@CalebChamberla6 Xometry went to crap when it became more B2B focused (among other things). These dinosaur companies are the heart of the problem you've fixed with innovation. Don't let them drag you down into their game of waste an inefficiency.