EXEPERTAI LAB
EXEPERT / DIRECTORY
← Blog

Slop Games Everywhere, but What About Coding? Maximilian Schwarzmüller on a New Model After the Hype

After the launch hype and a flood of AI-made 3D games, Maximilian Schwarzmüller finds the new model much better at 3D but only a good, not groundbreaking, coding model. What still matters: constraints, feedback loops, less over-engineering, and resisting hype fatigue.

Akmal Alif · 9 October 2026 MYT

A pile of glossy low-poly game objects spills out of a chat window on one side, while on the other a tidy code editor sits inside a frame of guard rails with a green test checkmark.

Every frontier model launch now follows a familiar script. Early-access posts call it a breakthrough. A few days later, once everyone can use it, the sober and disappointed posts arrive. Maximilian Schwarzmüller, a developer and course author who publishes a running commentary series called Code & Curiosity, watched that cycle play out while on holiday. The resulting video is a builder's view of what was left once the noise died down (Schwarzmüller, 2026).

The model in question is the one the video calls GPT-6 Astra, compared throughout with its predecessor, GPT-5.6 Sol. The model names and impressions below are the creator's. We have not benchmarked either model. The video is embedded below, and the timestamps jump to the matching moment.

Everybody's building AI slop games ... what about coding though? (Maximilian Schwarzmüller)The YouTube player loads in privacy-enhanced mode when you press play.Watch on YouTube ↗

The slop-game flood

The first thing everyone noticed was 3D. Schwarzmüller's timeline was full of Blender scenes and small games said to be made by the new model (watch from 0:30). Not all of it can be verified. Having tried the model, though, Schwarzmüller spotted recurring elements across the "vibed-up" games that suggest a shared source (watch from 1:00).

The guess is that the model was tuned heavily on 3D work during post-training. The jump in that one area is real and large compared with earlier models.

That fits what we saw from the other side this week, in Basic Dev's breakdown of AI game mashups. Models are especially good at producing things that already exist in quantity, and games are full of reusable parts.

Benchmarks are an entry ticket

Coding is what Schwarzmüller cares about most. The launch post shows strong coding benchmark scores (watch from 1:30). Schwarzmüller doesn't blame labs for optimising for benchmarks; a model that scored badly would never get tried, however good it was in practice. That incentive is exactly why the scores say less than they appear to (watch from 2:00). What counts is how the model behaves in day-to-day work.

Day to day: good, not a leap

For Schwarzmüller, the new model feels a lot like its predecessor for coding (watch from 2:30; 3:00):

  • A very good coding model. It has different quirks, but it is not groundbreakingly better.

  • Vibe coding still works. For internal tools and one-off projects, letting the model run with a loose prompt can get you far. Not every codebase needs to be perfect.

  • Maintained software needs discipline. If you plan to maintain and extend something, nothing has changed. You need constraints, you need to know what you are asking for, and you need to know how to work with the model (watch from 3:30).

It still produces ugly or needlessly complex code at times. One improvement stood out (watch from 4:00):

  • The older habit. The previous model treated every project like a ten-year-old enterprise system, piling on fallback paths and compatibility layers.

  • The new model. It does that less, but it still needs steering.

How far the baseline has moved

It is easy to forget how fast the floor has risen. A year ago, as Schwarzmüller remembers it, models stopped early, followed instructions poorly and wandered off. Now they can work alone on a task for a long time, provided the task is well defined and guardrails stop them drifting (watch from 4:30; 5:00).

The most useful advice in the video is about feedback. Give the model something it can evaluate or measure, and it will dig in and keep going until it solves the problem. That might be a test suite, a benchmark or a reproducible failure. The result may not be beautiful, but with a good setup it can be (watch from 5:00).

Hype fatigue is a real cost

The second half widens out. In Schwarzmüller's view, labs hype each release so hard that disappointment is guaranteed, presumably to please investors and keep subscriptions growing (watch from 5:30). Schwarzmüller compares it to blockbuster games that are over-promised before launch (watch from 6:00).

Combine that with new models almost weekly, and with real uncertainty about where programming and white-collar work are heading, and you get exhaustion. That mix is why model launches feel so frustrating (watch from 6:30; 7:00).

The verdict is plain. It is much better at 3D, and for coding it is "just a good model". Use it, and learn how to work with AI properly (watch from 7:30).

A builder's checklist for the next launch

Schwarzmüller's experience turns into a short routine for the next launch:

  • Test on your own work. Re-run a handful of your real tasks instead of trusting launch charts or early-access clips.

  • Watch for over-engineering. Compare how much code the model adds to solve the same problem. Extra layers are a cost you maintain.

  • Give it a way to check itself. Tests, type checks, benchmarks, screenshots: the more the model can verify, the further it gets on its own.

  • Write down your constraints. Architecture rules, naming, what not to touch. Steering is still your job.

  • Budget your attention. You do not need to switch models every week. Switch when your own tasks show a difference.

References

Schwarzmüller, M. (2026, September 9). Everybody's building AI slop games ... what about coding though? [Video]. YouTube. https://www.youtube.com/watch?v=8tfgjjRfSbY