Hey, quick intro π
Iβm the founder of NodeDB, currently running a SaaS + IT business. Most of what I do now revolves around building systems.
What I share here will mainly be things Iβm actually working on:
ML / AI framework and infrastructure, databases (NodeDB), edge to cloud distributed system.
Follow if youβre interested in these stuff. π
Around it: a tokenizer, a tensor library, a regex engine, an inference server, a notebook runtime.
30+ crates. Two years. One maintainer.
https://t.co/YldgKEBnfy
https://t.co/JliyAHpx3c
I scrapped this database twice before the third version worked.
NodeDB: multi-model, PostgreSQL-compatible, one binary. Vectors, graph, documents and time series in the same transaction. Embedded or clustered.
The first two attempts failed because I was patching a design instead of building one.
The third works because I finally understood why the first two did not.
I direct it, review it, reject it, and understand it down to the storage layer.
That part is not optional. I scrapped NodeDB twice. The third version works, and it only works because I understood why the first two did not.
I tried ELI5, ASD-STE100, and the recently released 'concise' style from Claude.
After a lot of experiment, this is my best output style.
It preserves all the details, while mantaining the quality.
Not the best at compression, but I like the response the most.
Try it:
https://t.co/QuzCtURNHg
Let me know your experience
Upcoming new Huggingface PipelineTokenizer is insane.
It performs better then gigatoken in certain workload.
I thought splintr already doing well, being competive with gigatoken.
Yet, another huge ceiling is coming.
@art_zucker good job π. Waiting for the official release.