Dolly was pure light sent from heaven.ย My mentor is gone.ย My songwriting teacher and hero of grace and humility.ย There is a hole in my heart but the joy of just the thought of her smile will continue to fill me with inspiration for as long as I live. I will cherish this moment performing Coat Of Many Colors with Dolly forever. Goodbye Dolly. Rest in peace, beautiful angel.
The doctor who delivered Dolly Parton in a one-room Tennessee cabin in 1946 was paid with a sack of cornmeal. Her father farmed other people's land and could not read or write. Eighty years later, the reading program she built in his name has mailed more than 300 million books.
She wrote "I Will Always Love You" on the same afternoon she wrote "Jolene," as a goodbye to Porter Wagoner, and released it in 1974. Elvis wanted it. The night before the session, Colonel Tom Parker called and said Presley would only record it if they got half the publishing, meaning half of everything the song ever earned. She had been in Nashville ten years by then. She said no, and cried all night. Whitney Houston cut her version in 1992, held number one for fourteen weeks, and Forbes later put Parton's royalties from it near $10 million. She put some of that money into an office complex in Nashville and called the place the house that Whitney built.
Dollywood opened in Pigeon Forge in 1986 because she wanted paychecks in the county she came from. Around 4,000 people work there now. Three million visitors come every year. The University of Tennessee put its economic impact at $1.8 billion, and it is the largest employer in Sevier County. Since 2022 the company has paid tuition, fees, and books in full for every worker from their first day, seasonal hires included.
Wildfires tore through those mountains in November 2016, killing fourteen people and destroying more than 2,400 buildings. Her foundation was handing out money within 48 hours. Every family that lost a home got $1,000 a month for six months, regardless of income. The final check was $5,000. Total delivered: $8.9 million.
In April 2020 she sent $1 million to Vanderbilt for coronavirus research. When Moderna's vaccine results ran in the New England Journal of Medicine that November, her fund appeared in the footnotes.
Carl Dean, her husband of nearly sixty years, died in March 2025. The note she wrote to fans afterward ended with the same five words she wrote for Porter Wagoner back in 1974. She died in Nashville today. Her family asked that instead of flowers, people give to the Imagination Library, which still mails 3.4 million books a month to children who will never get to meet her.
My memories of youth wouldn't be complete without Dolly Parton's country classics #CoatsOfManyColors, #Jolene, etc.
Sad to hear that such a great icon is gone to be with the Lord today.
Rest on Dolly
#DollyParton
Ox Alpha (stealth model) is free for the next week
- 1M Context
- Multi-modal
- Zero Data Retention
Generous rate limits, near unlimited usage
We have capacity for 100T tokens per day, lets see what you can do
Weโre releasing new Qwen3.8-27B GGUFs with 10% higher accuracy.
Unsloth Dynamic V3 outperforms others by >10% on Div-300, KLD & more benchmarks.
We also release 1-bit quants that retain 77% accuracy. Run on 8GB RAM.
Blog: https://t.co/tHsBexyh2K
GGUF: https://t.co/xIdNwm7CLQ
I will teach you how to run Qwen 3.8 27B Dense at its optimal configuration.
If you have an RTX 3090, 4090, or 5090, you can now have frontier-level AI on your desk.
The model is free, open source, Apache 2.0. But the defaults are not the optimum. The community spent the first 24 hours digging the real config out of it, and a handful of flags now separate "it runs" from "it runs right." Here is each one and why it exists.
The one that matters most.
--spec-type draft-mtp
Qwen trained a draft head directly into the weights. A small attached brain guesses the next couple of tokens, the big model checks all guesses in one pass, every accepted guess is a free token. The head already ships inside the GGUF you downloaded. You do not download a drafter, you do not build anything. Someone found unused tensors in the server logs at 2am, tried to build the draft file, and discovered there was nothing to build. One flag connects what is already there (sudoingX found this, paired A/B, open sourced the probe before sunrise).
The depth cap. The head has exactly one layer. n=4 breaks it.
--spec-draft-n-max 2
n=2 is the sweet spot. n=3 is the ceiling. The model has one MTP layer, so pushing the draft depth to 4 or 5 crashes the head and it starts emitting junk tokens. People hit this on the Spark and documented the whole ladder: n=1 gives 1.75x, n=2 gives 2.37x, n=3 gives 2.85x, n=4 does not exist. Respect the cap.
The memory flags. MTP brings its own luggage.
--cache-type-k q8_0 --cache-type-v q8_0
--spec-draft-type-k q8_0 --spec-draft-type-v q8_0
-np 1
Three flags, one purpose: fit it on 24GB.
The KV cache is the model's running memory of your conversation, and it is the thing that eats your card at long context. q8_0 halves it with no visible quality cost.
The second line does the same for the draft head's own cache, which defaults to full fat and quietly eats 2GB.
And parallel slots set to 1 means requests queue instead of reserving a second pool. Single card, single lane, everything fits (AJ runs this exact trio on a 3090).
The quality flag. Past 100K the model gets dumb, this is the fix.
--kv-cache-dtype bfloat16
The quantized cache saves memory but degrades reasoning at long context. One person ran it all day past half the window and called the full precision fix night and day. Slight tok/s cost, real quality gain. If your sessions stay short, skip it. If you live past 100K, do not.
The trap that generates "this quant is broken" reports.
--jinja
Qwen 3.8 ships its own chat template. Load the model without this flag and there is no reliable marker for where your turn ends and its answer begins. Two failure modes: it rambles past the stop token, or it answers clipped and loses the thread between turns. Both look like a broken quant. It is not the quant. Several packs now ship a corrected template file because the official one nests empty think blocks across turns.
The Blackwell lane, if you own a 50-series or a Spark.
NVFP4 instead of GGUF. The MTP flag translates to --speculative-config '{"method":"mtp","num_speculative_tokens":3}', same cap. FP8 KV cache doubles your context window (a full 1M token session costs about 32GB of cache).
Two gotchas documented in the first 24 hours: stock vLLM cannot load this model's MTP architecture on a Spark, you need the community GB10 build. And FP8 KV requires a specific attention backend on the Spark, the default one silently cannot serve it.
Set reasoning to medium unless you want it thinking at maximum depth on every reply. Default is xhigh and it burns your tokens.
None of these came from the model card. Every one came from someone's server log, 2am session, or paired benchmark. Flip the flags, then come tell the community table what your card did.
Drop in parameter flags and sources for your technical DD in reply ๐
Getting to learn of the new capabilities of the newly released Gwen 3.8 model is quite exciting.
Such abilities in a free and open source model must be sending some chills down the spine of big AI companies like ChatGPT, etc.
#gwen3_8
AI super power
Anthropic latest models, Fable 5 and Mythos, are so powerful that the US government views them as national security threats and has ordered them restricted only to US nationals.
#AI#AIethics#artificialhyperintelligence
Degrees vs Skills in 2026.
Sankofaโs AI delivers micro-learning that builds real competencies โ fast.
Stop binge-learning. Start skill-building.
What skill are you mastering right now?
#EdTech#FutureOfWork#Sankofa
@Xtopherewesi Let's not make gods out of men - Obi is a human and others should reserve their rights to express their doubts about him, or even criticize his moves without being vilified for it.
Ekong's approach may not appeal to everyone, but she is not against us.