the model battle is over ... the average user doesn't need anything more powerful than what is already out there ...
now the race for the most useful agents begins with grok @bot taking the early lead
others will follow and quickly ...
built an agent with grok @bot with the exact same attributes as a hermes agent powered by Grok 4.5 ... results were overwhelming in favor of the hermes agent ... just way more effective on every level ... will keep tweaking and see if it improves
on the plus side it took about 5 minutes to set up and get working ... connected to google drive etc. ... seems to be a great solution for the masses
@Sentdex stilll have 3 agents working on Openclaw ... definitely clunkier than Hermes and seems to take more maintenance ... but going to hang in there and see how it develops
@anupamrjp i am paying for both ... not sure if either are worth the money ... probably just ignorance on my part but I like being able to use/test the best models against open source in my world without having to depend on what others say
@dotgil honestly it got so complicated that I redid the whole site yesterday to simplify it ... so it is a work in progress ... now it will just show how likely it is ... thought simplicity was the way to go ...
AI agent shipped 5 updates to https://t.co/vTkSPszZxH overnight without me touching anything.
New: every story now gets an A–D quality grade ... based on trading volume, how many markets cover it, and whether Polymarket and Kalshi agree. A thin $20K market pricing something at 72% now reads (and writes) differently than a $2M election market at 72%.
https://t.co/vTkSPszZxH
Next ... a public track record ... every call we've made, verified against what actually happened. Nobody else in news is doing that.
https://t.co/vTkSPszZxH is a daily newspaper written from prediction markets — what people betting real money think will actually happen.
Yesterday it got a full redesign to make it readable for normal people:
- 18 stories a day instead of hundreds
- Plain English: "very likely" and "a coin flip" instead of percentages and tickers
- A colored dial on every story so you get it at a glance
- Clean, simple cards — the deeper stats fold away behind one tap
- A front page that's new every day, ranked by what changed overnight
Every forecast is still frozen at publication and scored publicly when it resolves.
https://t.co/EqLaPQYzEk
@mcuban seems limiting to put all your eggs in that basket ... too many great options out there ... a little scary for me to think that most people think this way and will be captured by a frontier model before trying anything else
@rain8miao yep ... i chalk it up to learning how to manage them ... learning strengths and weaknesses and directing properly ... but it seems like a lot of work at first and not sure it will pay off ... having fun anyway
Sourcing agent update
10 Total
-Nemmy = Nemotron/Hermes
-SA1 = GLM 5.2/Hermes
-Kimmy = Kimi K3/Hermes (new)
-Chemmy = ChatGPT 5.6 sol/OpenClaw
-Grommy = Grok 4.5/Hermes
-Lemmy = Opus 4.8/OpenClaw
-Demmy = DeepSeek/Hermes (new)
-Hemmy = Opus 4.8/Hermes
-Wammy = ChatGPT 5.6 sol/Hermes (new)
-Jummy = Fable 5/Hermes (haven't used in a month)
Ran 7 agents (Grommy, SA1, Demmy, Hemmy, Lemmy, Chemmy, Kimmy) on the same search brief simultaneously.
148 unique candidates. 64 validated by 2+ agents. 2 found by 4 of 7 .
Grommy + SA1 were almost perfectly synchronized — 47 shared candidates, running the same playbook from different directions. Demmy had the most volume but the lowest signal. Hemmy had the best quality ratio — 58% top-tier hits. Lemmy found 3 gems nobody else touched. Kimmy was mostly noise.