I'm done with Sol Ultra for now.
It's possible to make it work well in complex projects, but it messes up so much critical stuff and obliterates my weekly usage at the same time.
There were a lot of mishaps, but this was the nail in the coffin:
1. Asked it to write a major plan to refactor my market analysis system with 2 major goals:
- to make it faster (new features have made it slow)
- to make it scale in a RAM-friendly way when longer time horizons are analyzed
This is difficult because we're talking about hundreds of millions of DB rows.
2. After it used 35% of my Pro x20 weekly usage to write the plan, I asked a single Sol Max agent to implement it.
3. Sol Max finished its work and Sol Ultra was tasked with reviewing it against the plan. It pointed out several dozen P1 issues (like always) and Sol Max had to fix them.
4. After 4 of these review-fix improvement loops, Sol Ultra was fine with the results and told me the plan was 100% implemented. But it also told me that there was a risk I should know of:
The Sol Max agent had decided on its own to change the architecture described in the plan to a totally different one and implemented it.
This new architecture reduced parallelism by over 80%, which makes the entire system unusable.
I was shocked and asked Sol Ultra why it would tell me the plan was implemented when Sol Max violated the contract and implemented something entirely different, which didn't even contribute to our goals.
And it gave me the usual excuses... that it should have seen that. Yes, it should have seen that. It used 6 subagents concurrently to review the results and worked for hours on a single review, only to miss a huge plan violation in 3 of such reviews.
Instead, it told me at the very end when everything was supposed to be done.
That's inexcusable for a flagship model that drains usage limits this fast. It must care about well-written plans and goals, not ignore them.
IMO, Sol Ultra is in an experimental stage. That fits what I said in my previous posts. There are cases when it works really well, even better than anything else, but it's too difficult to handle and seemingly random, and its reliability is very low. It's too expensive a gamble for me atm.
@Sappples@CodeWithAmann For a long time people rode horses to get from point A to point B. Then we invented cars. Today, there are still lots of people who ride horses, but most of the time its for different reasons.
Writing code is the horse. Vibe coding with AI is the car.
@SahilPanhotra Vibe coding isn't bad. It's new. It's change. Before long, manually writing code will the be same as riding a horse as your primary means of transportation today. Sure, there will be people who do it, who love it. But won't be the norm.
@VadimStrizheus This is why I canceled my Claude Max sub and just went full on GPT buy in. Going on 2 weeks now, haven't looked back. I enjoy working in Codex better and while there are some things I'd opt for Fable over Sol, Sol is good enough.
@dieaud91 Yeah this last round of "frontier" level models really set the bar. Everything after only gets better. Exciting times to be alive and experiencing it in real time!