Behind the Games

How I Use AI in Yearbook — and What I Don’t Let It Decide

How AI helps word Yearbook clues while sources, facts, game design, QA, and final editorial judgment stay under human control.

AI is part of Yearbook.

That probably isn't surprising.

I built the game during a period when tools like ChatGPT, Claude and Codex were becoming capable enough that one person with an idea could suddenly build things that would have required a much bigger technical lift not that long ago.

Part of the reason I started working on Yearbook in the first place was because I wanted to learn what these tools could actually do.

Not in the abstract.

I wanted to use them.

Build something with them. Break things with them. Fix those things. Figure out which jobs they were good at and which jobs I regretted giving them.

Yearbook has been a pretty good laboratory for that.

It has also taught me one lesson repeatedly:

AI can be very useful without being the person in charge.

The earliest version was much more trusting

The simplest possible way to build something like Yearbook is pretty obvious.

Pick a year.

Send that year to an AI model.

Ask it to give you four interesting clues.

Done.

And, honestly, that can work surprisingly well sometimes.

A model can come back with four interesting facts, decent writing and a few clever clues in seconds.

The problem is the sometimes part.

I started seeing clues connected to the wrong year.

I saw facts that sounded plausible but weren't well supported.

The writing would get repetitive.

It would drift into a voice that didn't really sound like the game.

And sometimes the tone was just wrong.

A model might take a serious historical event and try to make it playful because somewhere in the instructions it had been told that Yearbook should be fun.

That isn't a behavior I wanted to discover after publishing the clue.

So the system gradually became less:

Hey AI, make today's game.

and more:

Here is a very specific job. Please do only this job.

That difference now drives almost the entire way Yearbook uses AI.

The facts come first

Before an AI model gets involved in writing the clue, Yearbook tries to establish what actually happened.

The current pipeline starts with Wikimedia's On This Day material and the articles connected to those events.

From there, the system works out things like:

  • the specific dated event;
  • what the answer should actually be;
  • which article best represents the event;
  • what supporting information is available;
  • which clue category it fits;
  • whether the subject is likely to be recognizable enough to use;
  • and how the event should be treated from a tone perspective.

That information gets locked into a record before the model writes the player-facing clue.

The model doesn't get to say:

You know what would be more interesting? Let's use this completely different event I vaguely remember.

It doesn't get to invent a URL.

It doesn't get to move something from 1994 to 1995 because that happens to make the clue work better.

It doesn't get to decide that the answer should be a broad historical topic when the dated event was really about a specific milestone.

The factual architecture is deliberately separated from the writing step.

That has made the clues much more reliable.

It has also made the AI's role much clearer.

AI is pretty good at the part where words go

Once the factual record is established, there is still a writing problem.

The source might basically say:

A specific event occurred on this date.

Which is accurate.

It is not necessarily a fun clue.

This is where AI is genuinely useful.

I can give the model the locked event, the evidence, the intended answer, the tone classification and the things it is not allowed to reveal.

Then I can ask it to turn that into a short Yearbook-style clue.

Maybe add a little wordplay.

Maybe offer a few possible kickers.

Make it concise.

Keep it grounded in the supplied material.

That's a very different task from asking the model to research history from memory.

And it plays much more to what I think these models are good at.

They are excellent language tools.

They can rephrase.

They can brainstorm.

They can generate several possible ways to make a clue less dry.

They can do that very quickly.

That saves me a ton of time.

I have no philosophical objection to that.

In the world we're in now, I would probably be making things harder on myself for no particular reason if I refused to use AI as a tool.

But a useful tool still needs a job description.

There are some things I don't want AI deciding

The biggest one is factual truth.

That should be obvious, but it is worth saying.

If Yearbook tells someone that an event happened in a particular year, I want that to be based on source material, not on whether a language model produced a convincing sentence.

I also don't want AI deciding what the game should be.

It didn't come up with the old-yearbook concept.

It didn't decide that the clues should get progressively more useful.

It didn't decide that close guesses should get partial credit.

It didn't decide that letter grades would be more fun than a generic point total.

It didn't sit in my family group chat and notice that people wanted to compare results.

It didn't hear from players who wanted an archive.

It didn't decide that the slightly crooked cards looked better than neatly aligned ones.

And it doesn't get final say on whether a clue feels right.

Those are product and editorial decisions.

Those are the parts I actually enjoy doing.

“Sounds like AI” is a real problem

This one is harder to measure.

Sometimes you read something and just know.

The writing is too polished in a very particular way.

Everything is a little too important.

A simple event becomes a “defining cultural moment.”

A normal change “reshapes the landscape forever.”

There is often fake enthusiasm.

There may be a grand introduction explaining why the subject matters before the subject has done anything to earn it.

And somewhere near the end, no matter what we were talking about, we arrive at a lesson about resilience, innovation or what it means to be human.

One supposed AI tell I’m not willing to surrender is the em dash. People seem to have collectively decided that using one means a robot wrote the sentence, which feels unfair to a perfectly useful piece of punctuation. Maybe AI uses too many of them. Maybe the rest of us weren’t using enough. I’m not an English teacher. I just like em dashes, and I intend to keep them.

Sometimes a trivia clue can just be a trivia clue.

I like puns.

I like dad jokes.

I like a clue that uses a slightly absurd comparison and makes me giggle.

I don't want every clue auditioning for a keynote speech.

So Yearbook has built up rules around writing style too.

It checks for repetitive or formulaic clue structures and flags wording that feels generic or awkward. The system also separates serious and sensitive subjects from lighter ones so the model isn't encouraged to treat everything with the same playful voice.

Those checks help.

But this is one of the areas where I still rely a lot on human review.

There isn't a perfect regular expression for:

This technically works, but it sounds like a robot trying very hard to be charming.

Maybe someday.

AI doesn't get to grade its own homework

After the model writes a clue, Yearbook runs deterministic checks on it.

That's an important word here: deterministic.

The system isn't just asking another model:

Hey, does this look good to you?

The clue is checked against specific rules.

Does it contain the answer?

Did a four-digit year sneak into the text?

Did it introduce a proper noun or number that isn't supported by the supplied evidence?

Does it fit the intended category?

Does the tone violate the rules for the event?

Is the clue malformed?

Did something go wrong with the player-facing wording?

If the answer is yes, the clue can be rejected or sent through a tightly bounded repair process.

The system can even handle certain answer-leak cases by replacing a complete revealed entity with a supported generic description, while still rerunning the checks afterward. More complicated failures go back for a rewrite instead of being guessed around.

This distinction matters to me.

AI can help create the wording.

It should not also be the only thing deciding whether its wording is acceptable.

That feels a little like letting a student write the exam, take the exam and then announce their own grade.

I am already giving out enough grades in this game.

Barbenheimer is a good example of the limit

One of my favorite examples of why the final human judgment still matters came from a clue about Barbenheimer.

The source event was about the simultaneous release of Barbie and Oppenheimer and the cultural phenomenon around it.

The automated process had enough information to connect the event to relevant material.

But the interpretation could still land somewhere around:

Oppenheimer

Which is not exactly wrong.

It's just not really the thing people remember.

The memory is Barbenheimer.

It's the double-feature jokes. The memes. Two completely different movies coming out the same day and turning into one weird shared cultural moment.

A few years from now, if someone says:

Remember that summer when everybody called it Barbenheimer?

that is the memory.

If the answer is just Oppenheimer, the clue has lost the interesting part.

So I added an editor-approved override that corrected the reveal while still keeping it tied to the original locked source material and one of the event's actual linked pages.

That is a very small thing in a free daily trivia game.

I also care about it a lot.

Both can be true.

The guardrails keep growing because people keep finding new edges

None of this was designed perfectly on day one.

The clue pipeline has changed repeatedly.

Players have emailed me when something was wrong.

Tests have exposed edge cases.

I have looked at outputs and thought:

Technically, yes. Absolutely not.

Each one becomes another little rule or process improvement.

Some clues had answer leaks.

Some used unsupported facts.

Some had category problems.

Some had awkward wording.

Some needed stronger tone protections.

Some source events pointed toward an article that was technically connected but editorially not the answer I wanted players to see.

The system now keeps much more information around during the draft and review process, including the underlying source record, QA results and editorial warnings. Rejected or incomplete days can stay rejected rather than being forced into publication because the calendar says a puzzle needs to exist.

That last part is important.

One of the easiest ways to make an automated system bad is to tell it that failure is not an acceptable outcome.

Sometimes the correct answer from the pipeline is:

No. This one isn't ready.

There is still a human at the end

The clue can pass all the checks and still make me change it.

Maybe the pun is bad.

Maybe it's technically accurate but confusing.

Maybe the answer is too obscure.

Maybe the clue is missing the culturally interesting part of the event.

Maybe I just don't like it.

That's okay.

I don't think automation should mean removing judgment.

For me, the appeal of using AI in a project like Yearbook is almost the opposite.

It handles enough of the tedious or technically difficult work that I can spend more time thinking about the parts I actually care about:

Is this fun?

Is it fair?

Does it sound like Yearbook?

Will someone recognize this?

Will somebody learn something interesting afterward?

Is this something I'm comfortable putting my name on?

AI helps me make the game.

It doesn't get to be the editor.

That job is still mine.