← Back to the dev blog

DEV BLOG / 04

A believable score can still be wrong

Following the first simulation results back to the assumptions that produced them.

Game screenshots
Theo Bishop’s game portrait with Idleball scouting, roster and all-pro emblems.
Player data and artwork from Idleball.

I had spent the previous stretch trying to get two teams through a game. Now I had results to look at, including ones I couldn’t explain particularly well.

Some of that was football. Upsets happen. Some of it was probably me. I needed a better way to tell the difference before I started “fixing” every score I didn’t like.

What changed since the last post

  • I went back through the rating assumptions. I looked more closely at how an individual player’s strengths and the rest of the roster were influencing a matchup.
  • Surprising results became something to trace. I tried to connect an outcome to the players and decisions behind it, rather than judging the score on its own.
  • I changed fewer things at once. Smaller adjustments gave me a better chance of understanding what had helped and what needed undoing.
  • I separated the questions I was asking. Whether a player profile made sense, whether the matchup used it properly, and whether the score looked plausible were different things to check.

Going back through the scores

An earlier results layout beside a more recent one, from different development saves.

Earlier results viewEarlier prototype
Earlier results screen showing the schedule and a recorded game in a development save.
The earlier results screen kept the schedule and a recorded result together.Open full screenshot
Recent results viewAndroid · Build 315 · Oct 4, 2026
Build 315 game-results screen with league scores and schedule paging.
The recent results view gives league scores their own rows. I can scan the results without reading a pile of output first.Open full screenshot

Learning to leave things alone

I had a habit of spotting three things I didn’t like and changing all three. Then the next run would look different and I wouldn’t know why. It felt productive while I was doing it. It was much less helpful when I had to work out what I’d actually improved.

I had created an excellent system for generating new mysteries. Unfortunately, I was trying to make a football game.

I got more disciplined about checking one part before moving on. That’s not a dramatic breakthrough, but it mattered. I was spending less time guessing which of my own changes had caused the latest problem.

There wasn’t one afternoon when the simulation suddenly became right. This was a lot of smaller adjustments, and a few ideas that didn’t survive a closer look. I was still learning how to judge my own work without either trusting it too quickly or assuming every surprise meant it was broken.

A single game could only tell me so much. Next I needed to keep the team and its results together across a season.

All blog posts