Behind the Games

How a Yearbook Puzzle Gets Made

A look at how Hindsight Games turns a year into four sourced, progressively helpful Yearbook clues—and why the process got more complicated over time.

I have always liked getting lost on Wikipedia.

You click on one thing you vaguely remember, which leads to another thing you definitely don’t remember, and 20 minutes later you’re reading about a Supreme Court case or a movie from 2002 or some scientific discovery that had absolutely nothing to do with why you opened Wikipedia in the first place.

I also like daily games. And old yearbooks.

Specifically, I’ve always loved those pages near the back of ’90s yearbooks that tried to capture everything that happened that year. They were usually colorful and a little chaotic, full of headlines and photos and blurbs about movies, music, news and whatever else somebody decided future teenagers would want to remember.

There’s something really warm about flipping through those pages years later. You see something completely ordinary — a movie poster, a song, a news story — and suddenly you remember where you were when it happened.

That combination eventually became Yearbook: four clues, one year.

The actual process for making those four clues has gotten considerably more complicated than I expected.

Which is probably my fault.

It started much simpler

The first playable version of Yearbook was actually pretty close to the game it is today. The yearbook look was there early: the colors, the slightly askew cards, the whole “page from an old yearbook” feeling.

The mechanics changed more.

At different points, the clue cards were stacked vertically. Sometimes you got more information at once. The scoring was less clear. I used colored squares and other systems that made sense to me because I had designed them, which is usually a warning sign.

People playing the game were basically asking: Wait, why did I get this score?

They also didn’t always like being treated as completely wrong when they actually knew roughly when something happened.

And that seemed fair.

If the answer is 1987 and you guessed 1988, you clearly knew something. That should feel different from confidently guessing 1962 four times.

That feedback led to the proximity system, where the game tells you whether you’re one year off, within five years, within ten years or nowhere close. The letter grades came later too, partly because they fit the Yearbook theme and partly because “B+” is just more fun than “correct on guess three.”

The proximity clues also created another little game inside the game.

If I have no idea when something happened, I’ll sometimes start with a round number — 1970, 1980, 1985 — and then use the feedback to narrow the possibilities. If one guess is within ten years and another is within five, you can start doing a little calendar math.

So even before getting into the trivia clues themselves, there’s some deduction involved.

First: pick the year

Everyone playing Yearbook on a given day gets the same year.

The year is selected deterministically from the date, which is just a technical way of saying the game doesn’t wake up in the morning, spin a giant wheel and hope for the best.

Once the year is known, the harder question starts:

What actually belongs in the puzzle?

A year contains an absurd number of things.

Some are instantly recognizable. Some are interesting but obscure. Some are historically important and terrible trivia. Some are great trivia and historically meaningless.

And some scientific discoveries may be extremely significant while also being something approximately eight people on Earth could identify from a clue.

That’s where the real curation starts.

Finding four things worth asking about

Yearbook draws from Wikimedia’s On This Day material and the articles connected to those events.

But I don’t want four random facts that happen to share a year.

Ideally, a puzzle has variety. Maybe there’s sports, entertainment, politics, technology, a famous birth or death, or something completely unexpected.

I also want the clues to become more helpful as you go.

Getting the year from clue one should feel impressive. By clue four, most players should at least have a reasonable shot at getting close.

That sounds easy until you try to define “recognizable.”

A person who grew up in the 1960s may immediately know something that means absolutely nothing to me. Meanwhile, I can see a clue about the first Tobey Maguire Spider-Man movie and instantly remember being in Mr. Parsons’ fifth-grade classroom when everybody was talking about it.

That movie did not alter the course of my life.

But the memory is incredibly specific.

That is a lot of what I want Yearbook to do.

So the current puzzle-generation process estimates how recognizable potential subjects are using things like Wikipedia popularity, how clearly the event connects to its article, whether it fits the clue category, and whether it can actually be described without immediately giving the answer away.

The system tries not to fill a puzzle with four deep cuts. Obscure things can be fun. Four obscure things in a row usually are not.

Then the facts get locked down

One of the biggest lessons from earlier versions of Yearbook was that simply asking an AI model to “write four clues for 1987” is not good enough.

It can produce something clever.

It can also confidently attach an event to the wrong year, repeat the same writing patterns, use a tone that feels weirdly inappropriate, or phrase everything in that unmistakably syrupy AI voice where every trivia question somehow becomes a profound reflection on the human experience.

Sometimes a trivia clue can just be a trivia clue.

So now the factual part comes first.

Before AI writes the player-facing wording, the system locks down the underlying event: the date, source description, intended answer, relevant article, supporting information, category and other metadata.

The model does not get to decide what happened or when it happened.

Its job is much narrower: help turn an already selected, sourced event into a concise clue.

That distinction has become really important to me.

AI is a very useful tool here. I would be making things considerably harder on myself if I refused to use it.

But the idea for the game, the mechanics, what makes a clue good, what the game should feel like, what counts as fair, the visual design and the final judgment about whether something should be published are still human decisions.

Specifically, mine.

Which means I also get to be responsible when something is weird.

Now try not to accidentally give away the answer

Once a clue is written, it has another problem:

It needs to contain enough information to be useful while somehow avoiding saying the thing it is describing.

This gets surprisingly annoying.

A clue can accidentally include part of the answer. It can mention a four-digit year. It can introduce a proper noun that isn’t supported by the source material. It can use a generic phrase so awkwardly that technically nothing is wrong, but no human being would ever write it that way.

So Yearbook now runs a series of checks before a clue can make it into a puzzle.

Among other things, those checks look for:

  • the answer appearing in the clue;
  • the year leaking into the wording;
  • factual details that aren’t supported by the source;
  • poor category matches;
  • awkward or repetitive clue structures;
  • inappropriate tone;
  • clues that are simply too obscure to be useful.

And the checks keep changing.

This is not some perfect system I designed once and proudly laminated.

People have emailed me when clues were wrong. Those reports led to better safeguards.

Someone once pointed out that the hidden year at the top of the game could be revealed by highlighting the supposedly obscured text with a mouse.

Obviously, you can cheat at Yearbook in about 400 easier ways. You can Google the clue.

But being able to expose the answer by dragging your cursor over it was still pretty silly, so I fixed it.

That kind of feedback has been incredibly useful.

There’s also a tone problem

The Yearbook aesthetic is intentionally playful.

The clues can be punny. I like dad jokes. I am a dad, so legally I may no longer have a choice.

But history is not always playful.

A system that happily makes a dumb pun about a movie premiere cannot apply exactly the same voice to a tragedy, disaster or other serious event.

So the game also classifies tone and applies stricter rules around sensitive subjects.

That part matters to me a lot.

This is a free trivia game, and I am probably more particular about a 25-word clue than a reasonable person needs to be.

But my name is attached to it.

If the game repeatedly gets facts wrong, treats serious material carelessly or feels unfair, people stop trusting it. And once you stop trusting a trivia game, there isn’t much reason to keep playing.

There are too many other things competing for two minutes of your day.

Then I actually look at it

After all of that, there is still a human review step.

Because a clue can satisfy every technical rule and still be bad.

It can be boring.

It can technically describe the right event while missing the part anybody actually remembers.

It can sound like AI.

It can just feel off.

One of my favorite examples of this is Barbenheimer, which probably deserves its own story.

The automated system found material associated with the simultaneous release of Barbie and Oppenheimer, but the resulting interpretation could essentially reduce the answer to Oppenheimer.

That is technically related.

It is also obviously not the memory.

The memory is Barbenheimer: that weird summer when two completely different movies opened together and the internet turned going to both of them into a cultural event.

A human understands that distinction pretty immediately.

The system needed help.

So it got an editorial override.

Finally, four little cards

After all that, what the player sees is… four cards.

That is probably my favorite part of this whole thing.

There’s sourcing, filtering, code, AI, validation, scoring and an unreasonable amount of tinkering happening behind the scenes.

But none of that should be the player's problem.

The player should see a clue and think:

Wait. I know this. What was that?

Maybe they get it immediately.

Maybe they use the year feedback to narrow things down.

Maybe they have absolutely no idea, finish the puzzle, flip the card over and click the Wikipedia link.

That last outcome is perfectly good too.

I do it all the time.

The point isn't for everyone to get an A+. Sometimes my family group chat produces an F and everybody gets to laugh about it.

You just probably shouldn't be getting an F every day.

That would be a pretty terrible game.