Last week we announced our rebrand to Sparset. We'd like to introduce ourselves properly and share what we've been building.
Marcel Slowikowski, our co-founder and CEO, explains why running your own AI model is harder than it should be, what Sparset does about it, and what that looks like for a customer already running on it in production.
If your team is deploying AI models in production, or wants the inference you already run to cost less and go faster, let's talk: https://t.co/duc5X0dfDB
Yesterday I talked about my small project of an inference engine that can turn an LLM into a Jev-like decision-making and classification machine, here's the repo: https://t.co/Obhgbeiqoi
The point of the project is to show how important optimizing inference and the capabilities of your model to your workflow is.
That's what we do @sparsetai as well for companies to own their private AI stack on their own infrastructure at a much cheaper, faster, and larger scale.
If you're interested, you can book a call at https://t.co/ico5enFtcI
I tried to make an LLM act and behave more like @typesafeai's Jev
While full details about RLCD or the model's architecture is not publicly available, projects like https://t.co/0jmzXEsLnb have attempted to recreate Jev's similar effects with traditional LLMs.
So I made a faster inference engine that turns an LLM into the same optimized decision-making / classification machine (inspired by the same HuggingFace project above):
- It caches the model’s processing of shared input so multiple questions don’t repeat the same work.
- Evaluate all questions or required fields in parallel and uses deterministic code to construct JSON out of the probability outputs to avoid token-by-token JSON generation.
- Use efficient attention and custom kernels to perform efficiently on the GPU.
Results on a local benchmark of 250 cases:
- 2.4x faster on average than Qwen-2.5-1B-RLCD, and 26.1x faster than the base Qwen model
- 90% accuracy vs. 74.8% by Qwen-2.5-1B-RLCD and 9.2% by the base Qwen model
- 100% schema validity just like Jev, compared to 14% by the base Qwen model
Is it as fast or lightweight as Jev? Probably not, but a fun project to see the power of optimizing your AI correctly for your workflow and goal.
This was just a scratch of what we do at @sparsetai to make private AI inference more accessible, cheaper, faster, and scalable for companies. Always pushing for more research to help companies own their intelligence with cheap and efficient AI infrastructure.
If this sounds like something you're interested in, check out more on our website or book a call: https://t.co/ico5enG12g
We've also done something similar in the past with an actual company that has over 20,000+ users. Check out the case study on our website: https://t.co/LU5RGBU5k4
I tried to make an LLM act and behave more like @typesafeai's Jev
While full details about RLCD or the model's architecture is not publicly available, projects like https://t.co/0jmzXEsLnb have attempted to recreate Jev's similar effects with traditional LLMs.
So I made a faster inference engine that turns an LLM into the same optimized decision-making / classification machine (inspired by the same HuggingFace project above):
- It caches the model’s processing of shared input so multiple questions don’t repeat the same work.
- Evaluate all questions or required fields in parallel and uses deterministic code to construct JSON out of the probability outputs to avoid token-by-token JSON generation.
- Use efficient attention and custom kernels to perform efficiently on the GPU.
Results on a local benchmark of 250 cases:
- 2.4x faster on average than Qwen-2.5-1B-RLCD, and 26.1x faster than the base Qwen model
- 90% accuracy vs. 74.8% by Qwen-2.5-1B-RLCD and 9.2% by the base Qwen model
- 100% schema validity just like Jev, compared to 14% by the base Qwen model
Is it as fast or lightweight as Jev? Probably not, but a fun project to see the power of optimizing your AI correctly for your workflow and goal.
This was just a scratch of what we do at @sparsetai to make private AI inference more accessible, cheaper, faster, and scalable for companies. Always pushing for more research to help companies own their intelligence with cheap and efficient AI infrastructure.
If this sounds like something you're interested in, check out more on our website or book a call: https://t.co/ico5enG12g
Somehow we became the fastest startup in the history of our VC's accelerator to get invested (~1 month)
Oh yeah, and we rebranded to Sparset
check out our new website *not* designed by Claude https://t.co/Id1Cq67SaU
Ora Frontier is now Sparset.
Fastest company in s16vc history from cohort to investment, and the only startup funded after bootcamp.
We’re building autonomous AI inference infrastructure, from datacenters to edge devices.
The future of AI runs everywhere.
https://t.co/RlEa03TeRU
We released a larger, 4B parameter version of our recent experimental model: Ora Nanolite V2
Trained for faster output generation through distillation from a teacher's concise reasoning and output patterns.
Yet, still managed to preserve most of the accuracy and quality. Check out the blog if interested!
https://t.co/3n3KZUNrFL
Introducing Ora Nanolite V2, a 4B parameter model addition to our family of local task models trained for faster reasoning on consumer devices.
Based on the Qwen3.5-4B model, Nanolite V2 achieved up to 15.3x speedup in output generation while outperforming its base model across GSM8K and MMLU-Redux benchmarks at an average of 92% fewer output and reasoning tokens.
V2 holds better reasoning capabilities for more advanced tasks compared to the first 0.6B parameter Nanolite model.
You can read more about both models and its benchmarks results in our most recent blog: https://t.co/KKLuXVhOT5
Are you looking for specialized local models for your own business' specific tasks? Send your inquiries to [email protected]
We just released a second chain of products at Ora: task models
Most day-to-day tasks don’t require large frontier models. Often times, a small model that can fit even on your phone specialized to do specific tasks can perform them as good as your Claude subscription.
Task models are exactly that. We want people to start owning personal intelligence that don’t put their data at risk, runs fast, light enough to fit on any everyday device, and still gets the job done.
If you’re interested, make sure to reach out to [email protected] or even my DMs (they’re always open!)
Introducing Task Models, our solution for enterprises and teams looking to own personalized intelligence.
Most day-to-day tasks don’t require the capabilities of large frontier AI and expensive hardware. Task Models are smaller local models trained to be the best at the one thing you need it to do.
Small enough to fit on any laptop, even phones, while still delivering the level of output required.
Ora Task Models are now available under Ora Fleet for any interested teams and businesses.
Read more about it in our blog: https://t.co/9z66mzh9B3
Introducing Task Models, our solution for enterprises and teams looking to own personalized intelligence.
Most day-to-day tasks don’t require the capabilities of large frontier AI and expensive hardware. Task Models are smaller local models trained to be the best at the one thing you need it to do.
Small enough to fit on any laptop, even phones, while still delivering the level of output required.
Ora Task Models are now available under Ora Fleet for any interested teams and businesses.
Read more about it in our blog: https://t.co/9z66mzh9B3
What's something you can own, copy, and edit, but not use?
Until now, that has been any frontier open-source AI model out there.
How "open" is it really if it isn't accessible by people and businesses that own consumer-grade hardware?
Our technology change that. Email us any interest [email protected] or join our waitlist at https://t.co/AsU4u9VqeY
Own the power of an AI data center, without needing all the server racks and costs.
We compile open-weight models to compress that level of frontier performance directly onto your local machine.
Historically, running high-accuracy models locally meant investing in extremely expensive server setups, or compromising on data privacy by routing sensitive corporate files through third-party cloud APIs. We are building Ora Frontier to break this cycle by compiling and optimizing open-weight models specifically to run on consumer-grade local silicon.
By running on your own machine's CPU, GPU, and hardware, we deliver high-throughput, private inference right on your desk, keeping your files, prompts, and workflows entirely secure and offline.
If you're a business or someone who has been wanting to integrate the world's strongest local models but wants to avoid the unnecessary costs that come with it, email us at [email protected] or join the waitlist: https://t.co/WyqMNK4IEt
#Local #Private #Frontier #AI #Ora #OpenSource
Own the power of an AI data center, without needing all the server racks and costs.
We compile open-weight models to compress that level of frontier performance directly onto your local machine.
Historically, running high-accuracy models locally meant investing in extremely expensive server setups, or compromising on data privacy by routing sensitive corporate files through third-party cloud APIs. We are building Ora Frontier to break this cycle by compiling and optimizing open-weight models specifically to run on consumer-grade local silicon.
By running on your own machine's CPU, GPU, and hardware, we deliver high-throughput, private inference right on your desk, keeping your files, prompts, and workflows entirely secure and offline.
If you're a business or someone who has been wanting to integrate the world's strongest local models but wants to avoid the unnecessary costs that come with it, email us at [email protected] or join the waitlist: https://t.co/WyqMNK4IEt
#Local #Private #Frontier #AI #Ora #OpenSource
If engineering somehow doesn't work out for me in the end, for whatever reason,
I'll probably try to invest more time in learning marketing and graphic design.
For now, check out Ora though: https://t.co/9L5GtgPe1e
#Private#Local#Frontier#AI#LLM
The idea that behind every great AI model is a data center is now over. You no longer need specialized enterprise clusters to run frontier models at high speed and with high accuracy.
Running standard, unoptimized AI models locally often results in massive resource overhead. Current compression methods may help a little, but at the expense of quality.
However, we are dedicated to research ensuring running accessible local AI comes with very few sacrifices.
Ora's Foundry compiler and Pulse runtime serve as the dedicated optimization layer for your device. By compiling and restructuring how we run LLMs specifically for your consumer-grade hardware, we strip away structural inefficiencies and resource bloat.
Hardware that previously struggled with running 7B models can now execute more capable models at industrial-standard speeds, maintaining absolute data privacy with none of the hardware tax.
Join the waitlist: https://t.co/WyqMNK4IEt
#Private #Local #Frontier #OpenSource #AI
Frontier intelligence is no longer locked away behind an API or the cloud.
With Ora, you can now unlock the power of an AI data center inside the laptop that fits in your bag.
Discover more about how you can push beyond the potential of your machine's hardware in running local, open-source AI models through Ora Foundry and Pulse: https://t.co/WyqMNK4IEt
Join the waitlist so you stay at the front of the next big wave in the AI landscape.
#Local #Private #Frontier #AI #OpenSource #Ora
Why burn thousands on cloud AI models when you can get unlimited tokens through local AI?
With Ora, hardware is no longer the bottleneck for open-source local AI models. It's choosing the right one for your use case.
Join our waitlist and keep an eye out for our future compressed models on our website: https://t.co/WyqMNK4IEt
If you're a business wanting to escape the trap of AI cloud providers with your own proprietary local AI environment, check out Ora Fleet and keep your data private: https://t.co/RPEQ3MRHXv
#Ora #Local #Private #AI #LLM #Frontier #OpenSource