Skip to content
asavschaeffer edited this page Mar 4, 2026 · 1 revision

Theory

The philosophical foundation for why events are the unit of meaning in game grammar.


Wittgenstein's Language-Games

The meaning of a word is its use in the language. — Wittgenstein, Philosophical Investigations §43

A videogame is a language-game. Consider a 2D Pac-Man-like game.

Inputs are discrete:

joystick ∈ {L, R, U, D}

State variables are discrete:

player.pos(x, y)
enemy.pos(x, y)
coin.pos(x, y)
berry.pos(x, y)

But state has no intrinsic meaning. Meaning arises only through events.


Collision-Defined Semantics

Game events are collision-defined:

player.pos touches wall        → movement blocked
player.pos touches coin.pos   → score increment, coin removed
player.pos touches enemy.pos  → game over
player.pos touches berry.pos  → buff gained

A coin is the thing that increments your score when you touch it. An enemy is the thing that kills you. These are not labels — they are behavioral definitions.


Conditional Rules

Rule semantics are conditional:

if buff active:
    player.pos touches enemy.pos → enemy removed

The same collision, different context, opposite outcome. Entities are defined solely by their behavior under collision.

The structure of what can follow what — which events are valid continuations of which histories — is the game's grammar.


What the Transformer Learns

A causal transformer trained to predict the next event token given prior context learns that grammar the same way a language model learns syntax. From this single objective it learns:

  • Physical regularities (movement, blocking)
  • Rule mappings (collision → outcome)
  • Temporal dependencies (buff duration)
  • Long-horizon behavioral patterns

Multi-head attention supports simultaneous modeling of mechanics, rules, and strategy without privileging any single description.


Connection to Event Sourcing

The connection to event sourcing is nearly 1:1:

Event Sourcing Game Grammar
Event log Token sequence
Projection function Transformer's learned prediction
Snapshots Periodic state snapshots in the sequence
Event schema Vocabulary design

Fowler identifies three capabilities that emerge when you guarantee all state changes are captured as events:

Complete Rebuild: Discard application state and rebuild it by re-running events. In Game Grammar, we do this constantly — the transformer learns by processing the event stream from empty states over and over. Every training step is effectively a rebuild from the episode log.

Temporal Query: Determine application state at any point in time. The hybrid tokenization approach (snapshots + deltas) is exactly this: snapshots give you a checkpoint, deltas let you roll forward. The model learns to answer "what happens next?" at any tick by implicitly reconstructing state from the event history.

Event Replay: When a past event was incorrect, reverse it and replay. This maps to how we handle invalid sequences during sampling — the validity checker identifies where the model diverged from game rules, and we can trace back through the event log to understand why.

The critical insight from Fowler is that the event log becomes the system of record. In traditional applications, this enables audit trails and debugging. Here, it enables interpretability: the transformer's "world model" is not a latent vector but a readable event history. You can inspect why the model predicted a collision by reading the events that established entity positions. You can debug a misprediction by replaying the sequence up to that point.

This also explains why snapshots matter: pure event sourcing is slow for queries because you must fold the entire log. Snapshots are cached application states — the model's "working memory" that lets it attend to recent state without reconstructing from episode start.

Clone this wiki locally