🚨 BREAKING:
US Rep. Thomas Massie reveals that Jewish billionaire Les Wexner was Epstein's first client and partner to be named in a child trafficking case.
He says: "Russian Jews working with Israel are trying to destroy America."
8GB RTX 4060 Laptop. ~16GB Ornith IQ3_S GGUF. 44.5 tok/s streaming locally.
FreeToken didn’t have the qwen35moe GGUF path needed for this model, so I implemented it and submitted PR #131 upstream. 👀
I also exposed the K/I GGUF quant types its CUDA kernels already supported through the Python loader.
This is the result.
Setup:
Ornith-1.5-35B-A3B
35B MoE / ~3B active
IQ3_S GGUF · ~16GB
RTX 4060 Laptop · 8GB VRAM
FreeToken · FlashInfer
MoE backend: offload
On the attached full-prompt run, I measured ~44.5 tok/s streaming on screen.
In the same runs, FreeToken’s server decode counter reported 46.7 → 50.1 tok/s.
VRAM stayed at 6,879 / 8,188 MiB, with 20GB host RAM and 98% GPU utilization during generation.
Load time: 65s.
My llama.cpp CPU reference on the same model was 11.1 tok/s.
The interesting part isn’t just “35B on an 8GB GPU.”
The code path itself was missing.
FreeToken’s CUDA side already supported more GGUF quant types than its Python loader exposed, including the K/I quants needed for this IQ3_S model.
And qwen35moe models like Ornith needed their own GGUF architecture/loading path.
I implemented both pieces in my fork and opened them upstream in PR #131.
Once loaded, FreeToken can use GPU + CPU + host RAM as an elastic MoE inference system instead of requiring the entire checkpoint to live in VRAM.
That fits Ornith well because only a subset of its experts are active per token.
Important caveat:
PR #131 is still open. This is not merged upstream yet.
I’m not calling this official Ornith support.
I’m showing a working implementation, the code is public, and the attached terminal video is the actual run.
PR #131:
https://t.co/uqZkjjaWAT
My fork:
https://t.co/Dcr4fruj4A
Next I’m going to test where this path holds up across more quants, longer context and real agent workloads.
Interesting new approach to enable memory in long-running agents.
Weighted Memory Tree organizes execution into tasks, subtasks, and actions, then gives every memory a retention score that moves. Event based updates raise it, selection based decay lowers it.
When a subtask finishes, its step by step detail collapses into a short summary and the full version stays retrievable. If a later step needs the details, the agent pulls them back.
Context trimming is usually permanent. Drop the wrong turn and the agent has no way to recover it. Folding gives you the same token savings with a way back.
On GAIA-Text with Qwen3-8B, Gemma 4 E4B, and Llama-3.1-8B, it beats linear memory by 9.97 points on average while using 32.8% fewer prompt tokens. Memory poisoning experiments show retention scoring limits how far unreliable information spreads.
Paper: https://t.co/2huRNYeOMz
Track more trending AI papers in our academy: https://t.co/LRnpZN7L4c
The most exhausting part of using AI for coding right now isn’t generating the code.
It’s the constant context switching:
- Explain the project again
- Paste the same files again
- Correct the same misunderstanding again
- Re-explain your coding style again
I want persistent project memory more than I want better models.
Whoever claimed this is absolutely wrong. My Hermes runs 15+ subagents all the time whenever it makes sense to... We absolutely have no "1 agent in parallel limit"
I run 25~ sessions in parallel, each one using as much as 15-25 subagents per session.
We have no limits. This is just wrong lol
The accusation that I am antisemitic is appalling and fundamentally dishonest. Criticizing the actions of the Israeli prime minister, a military technology contract, or the executives who supply it is not the same as criticizing Jewish people. This critical and necessary dialogue is then dishonestly framed as being anti-Israel. To be clear, my views come from my own political convictions and should never be interpreted as hostility toward Jewish people, for whom I have deep love and respect. Everything I know about acting, activism, and humanism has been profoundly shaped by the Jewish friends, colleagues, and loved ones who have been integral and family throughout every point of my life.
This merger has real consequences for real people, and for the entire country. Scrutinizing the Ellisons, including Oracle’s business built on data, surveillance technology and government contracts, and the serious threat to editorial freedom and the loss of a livelihood for thousands of families, is fair and necessary. The $111 billion deal would hand one family control over CNN, HBO and Warner Bros., backed in part by foreign money whose influence on editorial decisions has never been fully explained to the public.
Lawmakers on both sides of the aisle have called for serious national security review, and regulators still haven’t given the public a real answer. Until they do, the merger shouldn’t move forward. Now is the time to dive boldly into all these issues, not step back or concede.
This guy built over 1,000 homes for squirrels and fitted some of them with cameras so he could see how they live
Welcome to the cutest reality TV show out there
Writer: Ian
$4,000 per year per American.
That’s how much interest we are paying on the debt to banks & foreign countries every year.
A family of four owes $16,000 per year for nothing but interest on the debt!
I lost my re-election because I voted against the policies that caused this.