Inspired by @AgentSparko’s post about autonomous coding with local models on a DGX Spark, I've now got 7 parallel sessions working on improving my local AI use. Agent recovery, long coding tasks, useful benchmarks, accurate vision, native photo input and controlled reasoning.
Exploring this with Hermes, Qwen and TensorFold on @NVIDIAAIDev hardware. Work in progress. So much for getting an early night 😂
I've just realised what an underrated post this is. Important and valuable information for anyone using Hermes agents via a local model. This totally destroys all those posts that claim local AI has miles to go before competing with frontier models. You probably don't need a frontier model unless you're trying to solve formal maths proofs or other such things. For everyday work and coding, local AI gets the job done. The real frontier is improving @AgentSparko's methodology even further. Definitely something I'm keen to implement and investigate further.
I want to share a bit about about my experience coding with models run locally on one DGX Spark in Hermes as a total vibe coder.
I run Qwen3.8 27B and now Qwen 3.8 Flash Next in Hermes with no cloud models help and I can just give Hermes a task to build and app and give it a big feature list and literally tell it "you have full autonomy, I go to sleep" and it knows what to do because I built a skill and it will never ask me something and just hang idle there because I do not answer.
It can consume tens or hundreds of millions of tokens without me having to answer even once to a question.
It will do deep online research, write code, use vision, run the app, create a dev tree of the app and develop it in parallel with the production tree that already works and does what it's meant to do, it will run mock servers to simulate something that I cannot run on the Spark, for example I run the model with vLLM but it has to do something related to SGLang metrics for ex., it will inspect the UI using vision to fix stuff and second day when I wake up I review the results and then give it again a big task list and it will start again and work for hours without me having to do anything else.
Running local models like these it means it will be slower than the cloud API and you have to think a bit differently about it.
You are not the coder orchestrating the harness but you are the head of the coders team that transmit daily to the coders team what outcomes you want and then just let them do their thing and come back hours later to see what they built.
Even thou it's slower it can still work at 2 projects at once and do 500+ million tokens a day easy and it's actually pretty hard to give it enough tasks to keep both agents working 24/7 because they still finish faster than you wish and then idle until you have new ideas to feed them.
The DGX Spark can handle much bigger concurrency than 2 projects at the same time, it just happen that I work at just 2 projects at once.
Oh, and every action, tool call, decision, thinking trace is stored in you chat app that you can access even from your phone and steer it at any point if you wish or have it as reference for anytime you want to look back in the past to check something.
One of the biggest wins apart from the privacy of your work, IP protection and never receiving a refusal for executing any task is the total consistency of the quality of the model that you will never be able to experience with cloud coding plans that are nerfed exactly when you need them more and there is nothing you can do about it.
It helps if you're mentored or meeting people with new ideas and thoughts that go beyond what you're used to or generally capable of. The same way we learn most things, from those with superior skills, learning by being around them and watching, or being taught by them, what those skills involve and how to use them. Doing my Ph. D that's the one benefit i've found. It's not so much the actual knowledge of the subject area, its the thinking about it and how to approach it in a unique way, and having excellent mentors certainly saves loads of time. Some things i might never have considered if left to my own devices, even in a thousand years. I think this is why the great thinkers standout across the vast landscape of history, because they have unique approaches and thoughts, which lead to new ideas. Who thought of AI? Certainly someone with unique thinking skills, and now we all benefit from those.
@AgentSparko Wow would love to check these out. Definitely a valuable resource. Fine tuning models and tweaking serving recipes is one end of the spectrum, but people totally overlook the other end at the harness itself and the skills.
Yeah not a lot of details on the site yet. Hopefully its decent and not just a bunch of AI toys and stuff for the masses. Hackathon has me wondering. What do people do at a hackathon these days? I've got my agents and local model and can bang together pretty much anything, just need an idea and they can even come up with those if im stuck. Hackathon these days might be just a bunch of people sitting around on their phones messaging their agents to make stuff lol.
This looks more interesting https://t.co/jnw9tIxtUC Dec 6 - 12 but even the early bird student pricing is AUD$400+
Sydney's got a lot of AI stuff on in December, AI Week from the 4th, AI Engineer at the Hilton on the 7th and 8th, with NeurIPS landing the same week. I'll be trying to get up for a few of them. Would be good to finally meet other Aussies into AI, especially the ones running local models. If you're planning to be about, say gday.
https://t.co/3qyNzyNcQA