Lot's of talk about DeepSeeks latest models. Models that will be excellent for AI agent work.
It feels safe to say that if you want to keep ai token cost low, DeepSeek is certainly a good choice
For the ones of you that have followed me for some time, you will know that I have been a fan of Claude Haiku for general questions/chatting. It is just so cheap.
I ran quickly some models through my benchmark tool. The Deepseek flash model is 84% cheaper than Claude Haiku
Now, as you will see in the screenshot, the Deepseek flash model could be way cheaper if it would be more efficient with the output but I guess nobody cares with this cost.
[Experiment 7] I built a Tycoon game to learn about Hermes AI Agents
This experiment was tough mentally for me. I felt like I was going on autonomous mode myself and had difficult to find motivation 😂 .
This is probably why I feel I could improve the game quite some but I decided to just write up the article and move on to the next experiment. The need to conclude!
Anyway, the article is here:
https://t.co/L918FfYeQ6
Even though the game is not perfect, any player will learn a lot just by looking at the skill tree
The Hermes AI agent learning game (paperclips Tycoon type) got its overhaul from Opus 5 today and it definitely got better.
I am not sure why the Cursor Opus 5 agent was running for so long as the PR was completed after 20 minutes (it was running for 50 minutes before I stopped it), but anyway..
I asked it to review and improve the game and the skill tree. It didn't really change the skill tree but it added some dificult to unlock skills in each category. Those skills are only unlocked when a certain number of skills in the same category are unlocked.
See full skill tree in video.
As mentioned before, I find LLMs to be particularly good at creating skill trees for better understanding a subject like AI agents (in the end it is just graphs with nodes and edges)
I will play the game multiple times tomorrow to see if I learn something and note down all adjustments to be made
Quantization is interesting and highly relevant today when large models like Kimi is released. It is when we compress model weights so that we can run a model on a smaller computer / Datacenter
In experiment 4 game (linked below) we can invest resources into the following:
Quantization (INT8/INT4)
After training, weights can be compressed from 16 bits down to 8 or even 4 with little quality loss. The model becomes smaller and much cheaper to run — buyers pay more for models that are cheap to serve.
Quality loss is subject to each user and his/her personal evals but certainly it is a way forward for local AI
My next experiment, experiment 7, will be focusing on building purely a learn Hermes agent game. I think learning around AI agents is the most interesting pushing forward (until I change my mind🤣🤣)
I did already a simple AI Agent Tycoon in experiment 5 where the player could progress on LangChain, OpenClaw and Hermes but this time I will push the Hermes part much further
Similar to the other games I have built, kind of Tycoon style, I have given Hermes docs to the LLM to have a Cyberpunk 2077 style skill tree.
To progress on the skill tree generally provides motivation to learn more
The initial Hermes Skill tree comes up as follows:
- Setup
- Core Agent
- Tools
- Memory
- Messaging
- Media
- Automations
- Integrations
- Client Specializations (unique to the different customer types in the game)
Each one of the skill tree capabilities got something like 10-12 steps to acquire so game becomes quite massive
Screenshots below (still WIP)
My experiment 6, the autonomous drone defense game for learning about drones and autonomous tracking systems is now live
The game can be played for free here: https://t.co/HFKQIxKWaa
Substack write up will come later today or tomorrow
The game is quite complex and contains lots of ways to learn
The central part is that you as a player is running a drone alarm defense company. You earn skill points, reputation and money by completing contracts.
Skill points can be invested into a skill/capabilities tree.
Money can be invested into improving business aspects and the drone itself
Reputation increase payout and number of contracts for that type of customer
There are four types of customers: Private, Ultra rich, Industrial, and Government
Each one of the customers have their own demands so the player will have to decides where to invest skill points in the skill tree
The initial drone base is quite weak so it is not possible to put more than a camera on it but if a player decides to upgrade the base drone, advanced things like night sensors, lidar etc can be installed
I have mentioned it earlier but the skill tree contains the following areas:
- Hardware
- Flight Software
- Search and Navigation
- AI Vision
- Swarm Technology
So how does learning happen?
Learning happens because as a player you need to study the skill tree in order to invest build optimal solution.
As an example, in order to improve threshold confidence level for an object, you need to have human or animal classifier, which requires CNN object detection, which requires camera etc.
And the customer obviously requires that you detects only humans, and not a camera in itself
Learning also comes from the fact that total skill points is limited so you cannot just select all skills. You need to prioritize and optimize.
This is a concept I took from Cyberpunk 2077 which is annoying initially but later I realized that it makes want to replay the game with another build.
I personally learned a lot from this game and will as usual share more snippets over the following days.
Let me know what you think
[Experiment 4] - I built a game where I learn how an AI lab allocates resources in order to improve their models
It really is a so clever way to learn about things.
Here is the write up:
https://t.co/53fiQlIvHM
And the game is free to play here: https://t.co/Uh9UkVtxUk
Learn about a data centers by playing a free game!
Here: https://t.co/QnEHdBpQcE
The ones who knows me knows that I believe a lot in learning through playing games. I therefore asked Fable to build me this Paperclips inspired data center game so I can learn about computers etc
The game loop is easy, sell jobs to earn money, the more money you have the more you can upgrade your data center, and the more jobs you can sell.
Be careful with heating, bandwidth and power as those can disrupt your operations
No signups are needed and the game only takes a few minutes to understand
If you want to save your score saved and end up on the high score page, you will have to put your email though and you will be allowed to name the data center (the email is not shared with any one)
With Claude Fable 5 being out, I thought it would be nice to compare its pricing with Opus 4.8. Fable 5 is way more expensive
It can be seen on the screenshot that is in fact between 3 and 4 times more expensive on our MrBeast viral recommendation question
The reason being that it is using almost the double amount of output tokens 1676 instead of only 900 for Opus 4.8
Reduced the price now for StillPath app. Some impressions every day in the App Store but nothing is converting. Hopefully now it will lead to some downloads
I do believe in silent meditation and a progress path
#meditation for real
Still far away until I will reach Tokyo on WalkFar 😂.
At least now there is a clear distinction that I have left Lisbon. Still feels really motivating to take more steps like this though.
Can I reach Madrid in 2 weeks?
My 3 IOS apps are now available to download:
OceanBreath - breath hold practice : https://t.co/YDPLQuiIUS
WalkFar - every step you walk is applied to a virtual walk: https://t.co/si9136WwAi
StillPath - silent real meditation: https://t.co/fLohxQi1eg
Not enough people realize this, but if you ate a diet of only:
- butter
- tallow
- red meat
- eggs
- bacon
- fatty fish
You'd lose weight, feel great and be healthier than 90% of Americans