TradingView might be going out of business
The open-source TradingView Premium version is currently the #6 trending repository on GitHub, with 971 stars today
The devs also announced that the project and repo are going private
The version is available for Windows and Mac, with GitHub and Discord links in the "Community" section of the app
The true power of open-source
ICT ONE setup for life
This can get you funded!
Opening Range Gaps (ORG)
- 4:14 PM close
- 9:30 Am Open
What draw on liquidity to focus on
A complete model, Study!
People joke about how ICT traders don't trade specific days of the week/news events.
Well here is the actual reason as to why ICT doesn't trade Mondays:
CHINESE PROGRAMMER STACKED 4 NVIDIA CHIPS INTO A CLUSTER AND CONNECTED KIMI K2.6 - 300 AGENTS, 4,000 STEPS, $23,000/MONTH FROM ONE ROOM
4 NVIDIA compute modules stacked and cabled together - not a cloud server somewhere in a data center, a physical cluster sitting in his room processing AI workloads around the clock
Kimi K2.6 runs on top - 300 parallel agents working simultaneously, 4,000 coordinated steps per run - the kind of throughput that used to require a rented data center
before this: two weeks of manual research per project - TikTok, Reddit, App Store, competitors, reviews, one by one - now Kimi's agent swarm does it in 2 hours
$0.50 per million tokens via Kimi API - the cluster handles orchestration locally, the model handles reasoning - and the cost per research run is cents not thousands
research report, dataset, spreadsheet, dashboard, SEO report, competitive analysis - all outputs from one prompt into 300 agents
he replaced a research team of 20 people with one cluster and one model - and kept the margin for himself
$0.50 per request in - $23,000/month out
THIS GUY CONNECTED HIS AI AGENTS TO HIS OBSIDIAN AND BUILT A BRAIN THAT LEARNS ON ITS OWN. HERE'S HOW TO BUILD IT
Obsidian is just markdown files sitting in a folder. That turns out to be the perfect memory for an AI agent, because an agent can read and write those files directly. He wired his agents into the vault so they pull context from it, do the work, and write what they learned back. The notes aren't the point. The loop is, and it gets sharper every cycle
How to build it:
1. Point an agent at your vault. The fastest way, no plugins, no API keys: open a terminal and run npx obsidian-mcp /path/to/your/vault. That exposes your Obsidian folder to Claude as a tool it can read, search, and write to. Add it to your Claude Code or Cowork config and restart
2. Confirm it can see the brain. Ask it: "list the notes in my vault and summarize what's in them." If it reads them back, the connection is live. Now it starts every task with everything the vault already holds instead of from zero
3. Give each agent one job and a write-back rule. Tell it: "research this, then save what you found as a new note in /brain with links to related notes." One agent researches, one summarizes, one plans. Each writes its output back into the vault
4. Close the loop. Add one line to every agent's instructions: "read /brain before starting, write your result back when done." Now each task leaves the vault richer, and the next run reads that before it works. It compounds instead of resetting
5. You only steer. Review what the brain produces, point it at the next thing. The agents handle the reading, writing, and connecting
The edge isn't better notes. It's a brain that feeds itself, so the work gets sharper every cycle instead of starting over
Bookmark this
Chinese professor just revealed his development team - and it was 170 AI agents making every single company decision
not humans, not managers, not consultants charging $500 an hour
170 artificial developers working in parallel, never sleeping, never asking for a raise, never going on vacation
Kimi K2.6 runs all of them with one prompt and each one gets its own task
what used to take an entire department two weeks now takes two hours and costs less than a cup of coffee
and while most companies are still hiring people for these roles
the ones who understood what's happening already quietly rebuilt everything
this is what actually sits behind the growth of billion dollar companies right now
full breakdown of how it works in the article below
I've read a hundred "make money with AI" posts. Most are hype. This one is different.
The idea: build a bunch of boring little tools. PDF converters, GST calculators, JSON formatters, word counters.
Each one ranks for a keyword. Each visitor sees an ad. Do that 1,000 times and you're at $10K/month.
No audience. No content grind. No selling. Just tools that quietly rank and compound for years.
The best part is it's actually achievable. AI builds the tools now. The only real work is SEO and getting AdSense set up right.
That's exactly what I'm going to do next. Building the tools, ranking them, attaching AdSense.
If you haven't read this one yet, go read it.
And if you're not sure how to set up AdSense, I attached a video below.
One person rebuilt an entire company’s brain in 7 days inside Claude Code.
Not a doc. Not Obsidian. A living galaxy of nodes.
Every employee. Every AI agent. Every SOP. Every tool. All wired together on one screen.
Click a department. The human agent team opens up. The SOPs attached to it open up. What each person is allowed to touch opens up.
That last part is the whole game.
Permissions baked into the brain. An employee opens the chat, the AI already knows what they can access. Agents, data, SOPs surface inside the conversation like you tagged them by hand.
Obsidian can’t do this. Notion can’t do this.
No dev team. No funding round. No 6-month roadmap. 7 days, 1 person, 1 terminal.
This is the part nobody has priced in.
The tools to build $200K enterprise software now sit on your laptop for free.
The only thing missing is the guy who opens the terminal.
THE GUY WHO WON ANTHROPIC'S HACKATHON JUST GAVE AWAY HIS ENTIRE CLAUDE CODE PLAYBOOK FOR FREE. 10 MONTHS OF WORK, ALL PUBLIC
Affaan Mustafa won the Anthropic x Forum Ventures hackathon by building a full startup in 8 hours with Claude Code. Then he open-sourced the exact setup that did it. It's called Everything Claude Code, and it turns Claude from one assistant into an entire engineering team
Repo: affaan-m/ecc
This isn't a prompt pack. It's a system he refined over 10+ months of daily use shipping real products
What's inside:
A huge library of skills, dozens of specialized subagents, and ready-made commands, all working together. Each piece does one job. One subagent reviews security against OWASP standards. One optimizes memory so Claude stops forgetting earlier decisions around hour three. One learns from your past sessions and projects so the setup gets smarter the more you use it. Others handle planning, test-driven development, and language-specific code review
Instead of one assistant writing code, you get an orchestrated team. A main session delegates to the right specialist when the task calls for it, the way a real dev team splits work
The best part: it's not locked to one tool. It runs in Claude Code, Cursor, Codex and OpenCode, across Windows, Mac and Linux. Free, MIT licensed
This is the difference between using Claude like a search box and running it like a team that ships. The guy spent 10 months figuring out what actually works so you don't have to
Bookmark this
THIS GUY TURNED HIS OBSIDIAN INTO A JARVIS THAT TAKES A 3AM IDEA AND SHIPS IT AS A FINISHED PROJECT WHILE HE SLEEPS
The problem he solved: way more ideas than time to build them. So he wired Obsidian into a pipeline that takes a raw idea and carries it all the way to a finished project, with him stepping in only once
How it flows:
A 3am idea gets dumped into a single note. No structure, just the rambling
An automation reads that note and decides what it is. A project? A grocery item? A random thought? A TikTok to make? It sorts on its own
If it's a project, it moves to processing. The system researches it, watches the relevant YouTube videos, checks what tools already exist, and turns the mess into a proposed plan
Here's the only human step. He opens a Claude Code session and reviews the plan. Likes this, cuts that, approves it. That's the entire time he touches it
On approval the plan becomes a full requirements doc. Then one command, promote project, ships it to his machine and execution starts
A project manager agent spins up, reads the requirements, and creates the sub-agents that specific project needs. A website gets a developer agent. Research gets a research agent. They build it
Idea to execution, and he's in the loop for about two minutes
The trick isn't capturing ideas. Everyone has notes full of those. It's the layer that decides, plans, and executes without waiting on you
Bookmark this
a new 8GB VRAM GPU dense Local LLM leader was born yesterday
runs on: RTX 4060 / RTX 3070 / RTX 2080. any 8GB card
Qwen 3.5 9B (dense) was the go to for 6-8GB VRAM builds.
Gemma 4 12B QAT (dense) just changed that.
same llama.cpp + cuda 13.2. i7 12700H. 16GB RAM. same -ngl 99 flags. same 48k context.
unsloth gemma-4-12b-it-Q4_K_M.gguf
→ 15 tok/sec @ 48k ctx
unsloth gemma-4-12B-it-qat-UD-Q4_K_XL.gguf
→ 32 tok/sec @ 48k ctx
→ 26 tok/sec @ 64k ctx
64k context is a big deal. Hermes 3 agent requires 64k minimum to run. you're now getting full hermes compatible context on a budget consumer GPU at 26 tok/sec locally.
2.1x faster on identical hardware. and here's the part that breaks your brain:
the QAT-UD-Q4_K_XL is actually SMALLER than the Q4_K_M "XL" why?
QAT = Quantization Aware Training Google didn't train the model first and compress it later they trained it to be quantized from day one the weights already know how to survive low precision that's why you get more quality per byte
llamacpp flags: -m gemma-4-12B-it-qat-UD-Q4_K_XL.gguf -cnv -ngl 99 -c 48000 -v
fits in 8GB VRAM clean. no API. no cloud. no subscription.
and this isn't even the MTP variant yet
Gemma-4-E2B QAT runs on 3GB RAM, E4B on 5GB, 12B on 7GB, 26-A4B on 15GB and 31B on 18GB.
I have benchmarked the 26b and 31b qat as well on a single RTX 4090, checkout the comments for details.
If you have a 6GB or 8GB VRAM GPU, post your numbers.
more benchmarks and configs coming soon
NVIDIA just dropped Nemotron-3.5-ASR: one 0.6B model, 40+ languages, streaming.
parakeet.cpp already runs it. On a plain CPU, 2.5x faster than @NVIDIAAI 's Nemo runtime, output byte-for-byte identical (WER 0).
No GPU needed. Offline or real-time. Pick a language with --lang, or auto.
GPU numbers are coming to compare with Nemo framework.
IF YOU USE HERMES AGENT THIS IS THE GUIDE YOU HAVE BEEN WAITING FOR.
Nous Research just published the full official guide on how to use Plugins in Hermes.
Plugins are the layer most Hermes users never explore.
The ones who do describe the experience the same way every time:
They say the agent feels like a completely different tool.
10x more capable. 10x more useful.
Impossible to go back to running Hermes without them.
The full guide is at the link below.
Bookmark it before you close this tab.
Hermes Agent Desktop App is much more than a nicer UI, it's a full control surface for your agents.
I have my main Hermes Agent on my big PC with all my projects and skills and memories.
I want to do work on my laptop in the living room while watching baseball.
In less then 5 minutes, I'm all set up through the remote gateway. Watch to find out how, and look forward to my full guide on the Desktop app later in the week at https://t.co/QL8cYdHYrP
120 AI models. FREE for a year. no credit card.
Hermes Studio already has NVIDIA preset as the base_url - so all you need is a FREE API key and you're running 120+ models instantly
here's the setup:
1/ go to https://t.co/vKcoAoknWf
2/ register, log in, bind your phone number
3/ grab your API key
4/ drop it into Hermes Studio
what you get:
1/ 120+ models
2/ 40 requests per minute
3/ free for a full year
while everyone's paying for API access, this is sitting right there for free
Run Gemma 4 26B MoE on 8GB VRAM with 250k context at 20+ tokens/sec
If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware.
Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card.
The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?"
Today, I’m delivering exactly that.
I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!.
If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed.
The performance metrics are astonishing:
- 20 tokens/sec flat decode throughput.
- Stable, flat decode speed even with massive prompts.
- I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame.
# What about prefill?
Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable.
And this is running completely without Multi Token Prediction (MTP) active.
How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4.
The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse.
# The Test Setup:
CPU: Intel Core i7
RAM: 16GB System RAM
GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM)
# The Secret Sauce (The -cmoe Flag)
To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp.
This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache.
It prevents VRAM spillage and holds the throughput rock solid.
# The flags:
-m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v
Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking.
Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the replies