Building an AI-assisted visual story
A retrospective on shaping a short story for children with generated text, images, voice, and music—and the limits of the experiment.
I wanted to make a short visual story for children about two siblings finding their way through a forest. The idea was modest: combine a simple narrative with illustrations and audio, and see whether a set of generative tools could help me take the experiment from outline to a small, coherent experience.
Shape the story before the assets
The first problem was not image generation. It was deciding what the story was about and what a reader should be able to follow. I used a language model to explore a plot around siblings, animal characters, and cooperation. I revised those suggestions into a direction I could use as a basis for scenes.
I also tried to keep character and scene descriptions consistent when asking for images. A description can help make iterations more deliberate, but it does not guarantee that a generated character will look the same from one image to the next. That gap between intention and output became part of the work: review a result, adjust the prompt, and decide whether the variation was acceptable.
More than a text prompt
The original experiment combined a story draft, generated illustrations, synthetic voices, and background music. Each medium introduced its own review task. A line that works on a page may sound awkward aloud. An image may not match the scene. Audio adds timing and control questions that do not exist in a static story.
I did not establish that the result improved learning or had a measurable educational effect. The project was an exploration of making a small narrative, not an evaluation with children. I also cannot treat generated or sourced media as publishable merely because it appeared in the prototype. Image, voice, and music rights need to be checked before any of those assets are reused.
What I would do differently
I would begin with a smaller storyboard and a clear acceptance checklist for each scene: readable text, consistent character descriptions, useful image alternatives, and audio that can be paused or skipped. I would test the story with the intended audience only with appropriate consent and a clear way to evaluate the experience.
For now, this is a text-first retrospective. The original visuals and audio are intentionally absent while their ownership and reuse terms remain unresolved. A future interactive version would need keyboard operation, visible focus, text alternatives, and reduced-motion behavior before it could be considered complete.
Editor’s note
Originally published on 21 January 2024 and substantially rewritten in September 2026. This revision removes unsupported claims about educational impact and avoids presenting custom model behavior as a verified technical fact. The interactive story and its media are deferred pending rights and accessibility review.