As an update here, your PR absolutely crushed it! Drafting is upto 4.2x faster *on top* of my changes (so upto 140x faster than upstream llama.cpp overall). I have updated the article crediting you both on top and in a dedicated section at the end.
Also, you have fans on reddit :)
https://t.co/owc47UzTeA
@1a1n1d1y i've learnt that sqlite can scale up to 2 orders of magnitude more than people claim it does. it really does "just work" for a huge number of services.
@techNmak Amazing stuff man! I wish someone would do a big series of this for manifolds and tensors so I can finally understand GR beyond just how the equations work analytically.
As an update here, your PR absolutely crushed it! Drafting is upto 4.2x faster *on top* of my changes (so upto 140x faster than upstream llama.cpp overall). I have updated the article crediting you both on top and in a dedicated section at the end.
Also, you have fans on reddit :)
https://t.co/owc47UzTeA
I made drafting for prompt lookup decoding in llama.cpp up to 42x faster with four changes to its n-gram caches: 1) no map copies, 2) @sunitram's unordered_dense, 3) sorted vectors with a fixed-length binary search, 4) @lemire's constmap.
https://t.co/2k3sYw7Jxh
@valigo This is so cute! I want a cat on my desktop. What in Wayland's security model blocks this by the way? I am thinking it is some kind of permissions over the entire screen or something?
@LewisCTech If I use a framework is it still bespoke or does it need to be a hand styled css for each element? And if there isnβt even css, is it free range?
@KevinSzabo14 Happened to me two years ago. I was living in New York and my apartment had water damage. The ceiling on top of my desk just started leaking really bad after a storm