Suppose a generated quiz question has no correct option. The player gets it wrong and loses access to a spell. The problem reaches beyond the question: the game has applied a consequence the player could not avoid. When questions control resources or progress, test that whole path.

Where a question enters the rules

Johnson and colleagues describe Wizdom Run, where study notes become multiple-choice questions. Correct answers replenish mana; incorrect answers restrict actions. Their developer reflections report invalid options and difficulty mismatches (§§2–3). The study does not independently evaluate players.[1]

Check a complete turn

Choose one resource change in your prototype. Record the starting state, question, answer options, answer key, player choice and final state. Then trace which component accepts the question, decides correctness and commits the change. These are proposed checks for your own prototype.

Start with two questions. Can the game read the response? Is the answer actually supported by the input material? A valid JSON object answers only the first. Do not let a successful parse settle the second.

Proposed prototype checks; these have not been tested in the source game
Test inputWhat to inspect
Missing options or an answer index outside the listDoes the game reject the question before offering it to the player?
Readable output with no correct optionCan the player still receive a penalty? Who can review the question?
A prepared batch with unsupported difficulty labels or repeated answer positionsDo the questions match your stated difficulty rubric? Can position patterns become a shortcut?
A timeout or rejected question during combatWhat happens to resources, controls and progress while the game waits?

Decide what happens when a question fails

A replacement question, a retry and a pause each change the flow of play. Pick a fallback before the test and write down what the player should see. Inspect whether a delayed response can still change resources after that fallback has started.

For each case, compare the final state with the rule you wrote down. Keep the result even when the interface appears normal. If the answer was rejected, the resource change should follow the chosen rejection policy.

Sources and revisions

  1. Large Language Models in Game Development: Implications for Gameplay, Playability, and Player Experience

    paper · arXiv:2603.27896v1 (2026-03-29) · Source accessed:

Verification record

Paper reading
full_text
Code inspection
not_done
Execution
not_run

Prior full-text reading retained; official v1 HTML §§2–3 and §5 rechecked on 2026-10-10 UTC. No new figure, repository, game, model or QA execution. Proposed checks await human review.