Mad Scientist of Film, Video Games and Technology.
World First: Reinforcement Learning: Super Mario Bros. All Stars on physical console with cartridge!
For those who wonder what the Pro1X life is like.
When it is said and done, the issues in their compute, infra, and model output (quality, verbosity, and watermarks):
Avoid: The hype and delivery do not sync. Lost a month of productivity: hundreds of hours of effort, wasted.
@schwarzjn_@GaryMarcus As more companies realize they can make the targeted and focused solutions in house for better ROI then the current market, it's over.
https://t.co/N4KGMngOI0
AI is for the young, only they can understand it...
Oh really now?
In this episode: Old guy makes a local LLM platform for fun.
LLMs are not difficult, just big and painful like a dragon.
Train wreck is imminent:
They made promises to VC to get here, even more to the SEC to move forward, and their mini-skynet got loose... They're bleeding individual subs, and now company level. Fable fiasco, opus5 and sonnet5 suck. 'supervoting power' for founders since Dario only has an alleged 2% of the stock. They even have interview questions focused on what if stock goes to 0.
All while subjecting the world to eu law via: watermarks.
Without voting power, the IPO just adds day trade fodder.
How is this not already a train wreck. Perhaps I am just old, but I remember when companies wanted glowing good headlines before the IPO, not a dumpster fire, but hey modern finance...
@thsottiaux Why not add a toggle on the account to opt into 'all available' vs '5-hour rations'.
A check box is all it takes. The code already handles both modes obviously, so all it is now is exposing it.
@MTSlive "That's ok, I prefer to negotiate salary based on cash value only since stocks are always volatile and start ups tend to crumble in the transition from emerging to operating. VC is nice, but when it dries up and revenue matters, well with my services, you don't need to find out."
Wow, here I was just scrolling and get called out like that... 😳
- I hotwired my nintendo, so that I could teach python to play mario: my games are on cartridge.
- brilliant researcher, with a tower so old, the dust bunnies place bets on my gradient decent..
- occasional driver without a Ferrari, but a sub that hits limits in 0 to 60.
and I get my 4 sessions per day... at least when claude was worth it. AI is great, but never trust the thing fully.
Why do I need to spend more when the current crap can barely do what I want today?
@WesRoth Corp hates risk. That release - retract - rerelease, political focus over quality, and the last two models failed to deliver: all for increased cost, with excessive token bloat, and unreliable schedules.
remember, 11% of spend is only 5% of use, and has plateaued on adoption.
@GaryMarcus Talent can flee at any point, vc can dry up, their flagship product is a joke, support is pointless, watermarks, the fable fiasco, and I been with them for only 3 months.
the last one had issues nearly everyday, with costs passed to me in lost limits.
https://t.co/AHq1J21lnp
For those who wonder what the Pro1X life is like.
When it is said and done, the issues in their compute, infra, and model output (quality, verbosity, and watermarks):
Avoid: The hype and delivery do not sync. Lost a month of productivity: hundreds of hours of effort, wasted.
If @bcherny can use it and get it to work, cool: the only thing I would ask is if he can check his logs to see how much token bloat he gets hit with. All the excess talk, questionable 'mistakes', seems like everything is designed to eat tokens and drive away customers.
Claude has struggle sessions with itself on anything that is not boiler plate, and it has been consistent since 7/20.
I am working on challenges that most others walk away from, and Claude has been actively fighting against me.
I forked my working code base, tell it to take x, y, and z capabilities as is, explicitly telling it that any changes will break it: it is perfect as is, but still hoses it up.
Even when given explicit requirements, it does whatever it wants!
If they will not take the customers seriously, then we need to start with some class actions and law suits.
Lost a month of time, that's a lot of hours wasted and frustration caused.
@trq212
and still the company does nothing for all the delays and damages they cause?
FOR A MONTH NOW, IT HAS BEEN UTTER CRAP.
I don't want a reset, more crap is still crap. I want to see your execs start getting arrested for a long list of crimes.
How many tokens in limits are eaten with all the excess, all the defects, I lost whole sessions with the damn thing not connecting to it's own infra!
Give it instructions, it ignores them. I tell it the experiment is non-atomic, it adds gates to lock step to force atomicy! I can't trust the code i have because your failed product actively worked against me in the last 2 attempts to build the fork. A full month lost, and not even opus 5 max could get it's head out of it's ass!
If you can't be bothered to compensate for the impact, then as a company, you deserve the fullest extent of the law.
https://t.co/OhuAODgSbT
https://t.co/XsB2CqaIWL
https://t.co/Qr7uEANs0Q
First signs of the watermark in code confirmed.
Few things to note:
1- it added this comment to a var which was not in scope for modification.
2- it reordered the class so that this was near the top, this was far lower in the document.
3- ignores rules on comments clearly 'just ignored'.
I have my own test LLM setup so I can tell that is roughly 458 tokens for that one comment. collectively they add up, and I will get less usable amounts from the limits. I have no idea how much of this is elsewhere.
This is not opinion, it is not subjective: I had it do work, it consumed added tokens which will now add to the total token count of the files.
Before people cry: But it's only pennies!
It was also the plot of Gus Gorman in Superman 3.
@evermann@0xLuciferCore Perspective:
if you owed me or I owed you 54 dollars, then that is a large enough number; can get some pizza and 6-packs with that.
if either of us owed the other .34 of a dollar, neither would really care. perhaps on principle, how it was owed, but really?
World First: Reinforcement Learning: Super Mario Bros. All Stars on physical console with cartridge!
As per multiple searches, no one has ever connected a physical console to the computer and completed a Super Mario Bros.™ level.
There are plenty of examples with OpenAI's Gym code set, however those examples did not use a physical console, nor did they use a physical cartridge.
On November 11th 2025, episode 1034 completed Super Mario Bros. All-Stars™ version, level 1-1 on a physical console (Retron5) with a physical cartridge of Super Mario Bros. All Stars for the Super Nintendo.
While numerous searches for this type of experiment has not found any other examples of this, I wanted to document it so that we can say definitively:
Super Mario Bros. All-Stars™ version for the Super Nintendo Entertainment System™, has had level 1-1 cleared without memory access, using a physical cartridge with optical capabilities that emulate the human more then the game.
This was completed exclusively as if a human was playing the game.
In other words, It read the screen and pushed buttons.
@avrldotdev Planner should be passing an plan to the executor which should have a tool to validate it via structured code, returns a go/no-go w/ mods needed.
Sandbox containment via isolated hw dev env, and controlled input on tools.
Oh, and least privileged access, it's a thing!
@jumperz How many of these power users and pro 1x accounts, make the decisions for those larger companies?
As a product manager, if you tell me half my bill is errors and bloat, no way! My job is to make sure my team has all the right tools and means reliability and trust in my vendor.
9 symptoms, 1 root cause: token bloat, where you add as much verbosity to increase token count.
The unspoken process:
First you enter your prompt, costs you input tokens but you can control that.
Really? did you forget you have your question, and some skills, and Claude will evaluate skills for you from it's collection, so you thought you used 200 tokens, but the model saw 2000.
then "it's thinking" eats tokens, then the guard rails, and the watermark, then output tokens. agents is just a fancy way to say we took the output and sent it through the cycle a few times, the token burn masterpiece.
When the next prompt or use of those files, even more input tokens in context.
And that is the magic:
Those new tokens of waste will create more tokens of waste when you ask the very next prompt.
the best part is the more they do that, the quicker the limits are used, the faster you need to give more money. wouldn't want to let an inpatient species like humans wait 5 whole hours. Don't want your work slowed down, do you?
@TokenGremlin extended usage? you mean for the low low price of letting them own one of your systems to run their code, they will let you use your entire week of limit or is it an extra %50 of the %50 you already had so %75 of your limit?
I canceled because they suck:
https://t.co/2rTVHiEBk9
For everyone who wanted to know what it is like to have Claude AI Assistance. This is what it actually looks like when not running agent harnesses.
I would advise parsing those logs ASAP!
If you're wondering why the API bill is up there with no results, they you might have had a week like mine!