How game development led me to synthetic customers
Building small games led to a testing concept now in development: synthetic customers that exercise software through realistic scenarios and expose behavior over time.
I have been spending a lot of tokens building small video games.
Some are for my kids. Some are for me. They are simple satisfaction projects: make something, play it for a while, enjoy what works, then move on. From the outside, this looks like a fairly disposable use of capable agents. The game exists because it is fun to make and fun to play. It does not need a business case.
But the games kept pushing me toward a problem that turned out to be useful well beyond games.
Getting an agent to build a playable version is only the beginning. A game can run without being good. The controls can work while still feeling awkward. A level can be complete while being too easy or too hard. A strategy game can look balanced for the first few minutes and fall apart once its economy has been running for much longer.
The code tells me whether the game works as software. Playing tells me whether it works as a game.
Once that distinction became obvious, I started thinking less about how to generate another feature and more about how to build the feedback loop around the thing the agents had already made.
A game needs a player, not only a builder
Small iteration cycles help, but they create a new job. Every change needs someone to experience it.
Is the game too difficult? Are the controls usable? Does the level design make sense? If the game has an economy, what happens after enough time passes for the early decisions to compound?
I can answer some of those questions by playing. My kids can answer some of them too. But human play is slow, especially when the interesting effect appears later. Repeating the same route or waiting for a strategy system to develop takes time. A small change can restart that wait.
So I began assigning agents to play the games themselves.
The role changed. An agent was no longer only producing the game. It was entering the game, trying it, and feeding another iteration. In a strategy game, the simulation could run faster so the agent could observe longer effects without waiting through the same amount of real time. That made it possible to think about difficulty, controls, level design, and economics as parts of a repeatable loop rather than occasional human checks.
This is still imperfect. An agent playing a game does not magically decide whether the game is fun. The useful part is narrower: it can repeat scenarios, move through the system, and create feedback at a scale that changes what I can inspect.
That narrower capability was enough to unlock the next idea.
The game stopped being the interesting part
At some point I looked at the game loop and stopped seeing it as a technique specific to games.
The agent had a world, a set of available actions, and a scenario to follow. Its behavior created effects inside the system. Some effects appeared immediately. Others became visible only after more simulated time.
That pattern exists in business software too.
A real customer does not experience a product as a collection of isolated features. They use it repeatedly. One action changes the state for the next one. A workflow that looks correct in a short test can feel very different after it becomes part of daily use. New behavior can have effects that only become visible after time and repetition.
The game work made me ask a different testing question: what if agents could play the software from the perspective of a customer?
That question shaped a concept we are now implementing at appointmed. We call it synthetic customers.
The basic idea is deliberately simple. An agent uses the features available in the product through a realistic scenario. It behaves like a customer using the software over time, not like a test that calls one function and checks one response. The same kind of loop that lets an agent play a game can let it exercise a product, repeat a workflow, and help us observe how the experience changes as the scenario develops.
Time is part of the behavior
The most important connection came from strategy games.
Some systems cannot be understood from one move. Their behavior emerges after many moves interact. An economy that looks healthy at the start may drift. A decision that appears harmless may become dominant once it is repeated. To see that, I need either patience or a way to make time move faster.
Game simulations make the second option natural. Fast forwarding is part of how I can inspect a longer arc without living through every second of it.
The synthetic customer concept carries that thought into product testing. If a scenario represents repeated product use, simulated time can help expose longer effects that are difficult to reproduce through ordinary human testing. The point is not simply to perform more clicks. It is to see what repeated use does to the scenario.
That changes the unit of testing. Instead of asking only whether one feature behaves correctly at one moment, I can ask what happens when the feature becomes part of a continuing routine.
This is the ambition. It is not a result I can claim yet.
The concept is still in the making. We have not reached the point where I can tell you that synthetic customers caught a particular problem, predicted a real customer's experience, or changed a release decision. Publishing those claims now would turn a useful work in progress into a story that is tidier than reality.
What I can say is that the game experiments changed how I framed the problem, and that framing became concrete enough to shape a feature we are building.
Build the feedback loop around the artifact
The practical lesson is not that every team should create synthetic customers. It is that agent work becomes more interesting when the agent can interact with what it produced.
The first question is usually whether an agent can build the artifact. Can it write the feature, construct the level, or change the system? That question matters, but it ends too early.
The next question is whether an agent can inhabit the result. Can it move through the available actions, repeat a scenario, and reveal behavior that is hard to see from the implementation alone?
For games, that means playing. For software, it may mean following the product the way a customer would use it over time. The exact simulation depends on the system, but the operating rule is the same: do not stop the loop when the code runs. Keep going until something exercises the behavior you actually care about.
That also creates a healthier boundary for what the agent can and cannot tell me. It can provide repeated observation. It can help surface patterns. It does not turn a simulation into a real customer, and it does not remove the need for human judgment about what the product should become.
I started the game projects because I wanted small things my kids and I could play. I did not expect them to influence how I think about testing a healthcare SaaS.
That is why I keep making room for these apparently throwaway builds. Sometimes the artifact itself is temporary, but the feedback loop it forces me to invent survives. In this case, agents that learned to play small games gave me the shape of a much more serious idea: software exercised over time by synthetic customers, while the real work of understanding the results remains with us.
Get the build log