I'm getting many questions about a small plugin that I've built in the Hermes desktop app, which shows the actual generation tok/s. Not a benchmark, not a false figure โ just realistic work with real tasks.
You can easily spot how this number compares to the promises from the recipe.
Happy to see this on your screen?
@Teknium
Starting September 14, we're permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will be in place.
@OwlCoreAI@MiaAI_lab no mate, this looks amazing! Should work for you really fine :) you will see when is slowing down and when is speeding. Better thing than benchmarks
@MiaAI_lab I build this with GLM-5.3 Flash - just asked to build plugin and add this to Hermes bottom right window.
Is showing in real time tokens generation and total token generated for current session
@t0nil0@matteocollina@NVIDIAAI tbh all depends from model and recipe = CPU and GPU usage. When I'm using some small models like GLM-5.3 flash or DeepSeekV4 flash fans are stable - literally no noise. Full GLM-5.2 or GLM-5.3 (dint' try this yet) noise can be pretty high while processing prompts.
@bridgemindai You are blindedโฆ you are talking about here and now. Wait till December this year and tou posts will change dramatically.
While people like you are complainingโฆ other ones already well prepared with equipment for running locally big models.
Trust me