Document Type
Article
Publication Title
Applied Sciences-Basel
Abstract
Although Large Language Models (LLMs) produce fluent prose, they routinely tend to generate narratively sterile content for interactive media—backstories in which every conflict has been resolved before play begins, giving rise to what we call the Closure Paradox. Moreover, current evaluation metrics—from n-gram overlap to retrieval-augmented faithfulness scores—inadvertently reward this exact type of closure in cases in which it is undesired, leaving game designers without tools to measure whether a generated character is actually useful as a foundation for play. We propose a dual-axis evaluation framework that separates Temporal Consistency from Ludic Potential. Two new metrics anchor the framework: the Ludic Potential Index (LPI), a weighted measure of a backstory’s openness to future play, and the Potential Conflict Value (PCV), a per-event measure of playability fuel. The framework is implemented as a three-agent pipeline—Notary, Judge, and Stress-Tester—whose judge layers are stated in G-Eval terms so that they can be driven from standard open-source evaluation harnesses such as DeepEval and TruLens, and is validated through synthetic stress testing in which a separate LLM agent attempts to instantiate a Session 1 encounter from the generated backstory. The paper includes a comparative analysis of a correction mechanism implemented during generation (gated generation—where explicit metrics are applied to filter or constrain outputs) versus direct generation in the context of both commodity language models (e.g., Qwen or Llama) and frontier versions (e.g., Opus). The methodology evaluates the framework interventionally using a semantic, judge-based Adaptation Effort formulation (a metric of the additiona work required to make new content playable for the character). The results show that (1) applying the gating mechanism yields a meaningful, medium-sized reduction in this Adaptation Effort (Ea), successfully reining in variability; (2) however, this benefit in frontier models produces only a marginal change, showing stability because these models naturally produce low-Ea drafts from the outset; (3) small-size commodity models do not benefit from the gate mechanism either, because their limited context handling does not profit from the iterative rewriting; and (4) nonetheless, the combined effect of the gating mechanism and the generator tier reveals a directional trend. The conclusion is that gating affordable commodity generators with explicit metrics meaningfully reduces downstream narrative preparation load; moreover, the presented set of metrics articulates criteria for selecting models and correction mechanisms, and their use can be applied to other narrative quality control processes. We argue that strategic vagueness is a functional requirement, not a defect, of interactive narrative—and that measuring it formalises the neuro-symbolic bridge between LLM generation and the symbolic state-tracking that interactive storytelling has required since TALE-SPIN.
First Page
1
Last Page
33
Publication Date
9-2026
Language
eng
Recommended Citation
Peña, L., Peña, J. M., & Buren, R. (2026). Toward Ludic-Aware Narrative Generation: A Neuro-Symbolic Framework for Evaluating Playability in LLM-Generated Backstories. Applied Sciences, pages 1-33.
