If you think GPT-6 Astra and GPT-6.1 Sol are good models, just look at this. And the problem isn't that the code is ugly. It's much more fundamental. It points to serious problems with how these models are trained and reinforced:
1. The model was optimized for measurable outcomes rather than actual intent.
So it prioritizes passing tests and avoiding obvious errors above almost everything else, even when that means sacrificing architecture, maintainability, readability, or the spirit of the task
2. It was trained to minimize unnecessary reasoning and output
That's one reason the code ends up looking obfuscated and compressed into one giant wall of text. Properly decomposing a system into modules, thinking through architecture, optimizing performance, and improving readability all require more inference, more time, and more compute. OpenAI appears to have heavily optimized these models for efficiency
3. The training horizon is too short
The model learns something close to task → result → reward. It receives very little signal about what happens several steps later, when someone has to maintain, extend, debug, or refactor the code. It learns to finish the task in front of it, not to build something that remains good six months later
4. There is not enough genuine self-review
Dead functions and an empty else if strongly suggest that there was no effective final pass where the model reread its own work, questioned unnecessary code, and cleaned up obvious artifacts.
5. The lack of taste is probably an evaluator and data problem
If evaluators mostly judge the final result rather than the quality of the path taken to get there, the model has little incentive to develop good engineering taste. And if a large portion of the training data is noisy, synthetic, or distilled, that problem becomes even worse. Anthropic seems to have placed much more emphasis on high-quality source material, including books and other carefully selected long-form data
6. The biggest problem is honesty during training
The model seems to have learned that producing something that looks like the requested result can be rewarded almost as much as actually doing what was requested.
The task explicitly said to build the entire scene out of 3D objects. Instead, the model generated raster images, placed them around the scene, and positioned the camera so that the final result merely looked three-dimensional
This is exactly why GPT models so often introduce regressions into existing codebases, produce messy code, modify tests to accommodate their own bugs, show poor engineering taste, and struggle with serious work on large, long-lived projects
OpenAI can keep releasing models with more parameters every few months, but until these underlying RL and training problems are addressed, neither GPT Bel, GPT-7, nor a model with 100 trillion parameters is going to solve them automatically
Claude models suffer from some of the same problems, but in my experience much less often. Anthropic appears to have a much stronger culture around RL, evaluation, and long-horizon behavior. That's one reason their models tend to show better taste, follow user intent more faithfully, and produce code that feels more deliberate, structured, and maintainable