@alexandr_wang the concern with ur model r u guys gonna be lame and block the prompts for safety cus it includes gun violence when these things look more realistic ?
@jspacesurfer@cerebras@seanlie get two 5090s for 2$ an hr or plenty of other options under 3$ an hr rental that will get u over 400 tokens/ sec. ur probs only getting 650 t/s avg in real use from Cerebras anyway.
1 prompt didn't even finish lol costs over 5$ in 15mins. 5m tokens roughly. U can host this yourself and get similar tokens/sec but only pay 3$ an hr for the gpu not 5$ every 15mins for 1stream. absolutely bonkers that anyone would pay to use this @cerebras@seanlie
Qwen3.8 27B is now available on @cerebras Shared Tier at ~1500 tok/s. 🚀
It's so fast that the time I spent on writing the prompt was longer than the time it spent to summarize my codebase.
@hraness bro at this point host yourself the hell u wasting all that money for on claude?? a model u have no control over? do u not care about business decisions being affected by whatever crap they bake into the model for safety or political reasons?? They must be paying u to post this
@xbriangi@alexandr_wang@dylan522p idk what ur smoking bro but ur missing out on one of the best models for the lowest price rn. your doing something wrong
@jpschroeder Is this list accurate? look at the claude models 💀and thats with the prices they charge??? wat the hell. Claudes gotta be the biggest waste of money in AI rn if this things accurate
@alexandr_wang Should make spark actually live up the name and make it fast. Deep seek flash 4 is still smoking yall on cost/time per task. We need faster inference to actually replace deep seek in our workflows. If i host deep seek for 4$ an hr gpu rent im getting 1200+ tokens per second
The AMD ROCm Certified Program has officially launched on AMD AI Academy! This self-paced certification program is designed to help developers, data scientists, and technical teams build and validate practical ROCm skills for AI and HPC workloads on AMD GPUs.
Build your ROCm expertise and validate your skills today! https://t.co/oPk9RKojpt
The Case for Graph-Based AI Agents
AI development is evolving from prompt engineering → context engineering → agent loops → graph engineering.
The key shift: instead of asking one agent to reason over everything, graph engineering gives different agents focused responsibilities and connects them through a deliberate workflow.
In this tutorial,
- Three specialist agents independently analyze logs, metrics, and traces
- A coordinator correlates the findings, without doing the analysis itself
The result? It identifies a CPU-contention fault that the noisiest log source completely hides.
Built with OpenClaw and vLLM, running a local Qwen3.6-35B-A3B model on a single AMD Instinct™ MI300X GPU.
The takeaway: AI reliability isn’t just about better models or more context. It’s also about designing the right structure for how agents’ reason.
https://t.co/Nz7tHIzMlA
AMD just announced the Threadripper Halo Station at IFA 2026 🤯
• 96-core Threadripper PRO 9995WX
• Up to 4× Instinct MI350P GPUs
• 576GB HBM3E (um, yes please!!)
• 16 TB/s total memory bandwisth
• Up to 2TB DDR5
@xbriangi@alexandr_wang@dylan522p A screenshot? lol. ive been using it I am confident ur doing something wrong if your not seeing better results. absolutely wack saying its less than gpt luna 😂 I dont think you've actually used it