How to Make Manga with AI: The 5-Stage Workflow from Script to Finished Page

Start with the honest part: no tool today takes a story in and gives a finished manga out. That does not mean AI is useless here. Split manga production into five stages—script, composition, line art, finishing, lettering—and the middle three are where AI pays off most. The two on the ends are storytelling decisions, and handing those over usually makes you slower.

What actually stalls most people is not image quality. It's the handoffs: how to write a script a model can read, what a character needs before you generate a single panel, and whether to refine or start over when a result is wrong. This article walks all five stages and the handoffs between them.

Why "one click, one whole chapter" doesn't work yet

This isn't a compute problem. Three things manga requires sit outside what image models currently do well:

  • Consistency across panels. The same character has to keep one face, one outfit, and one height ratio across dozens of panels. Every generation is an independent event—"remembering the last panel" comes from reference images you supply, not from the model.
  • Text inside the image. Speech balloons and sound effects are part of the layout. Generated lettering tends to come out as garbled or hallucinated text, and it's worse in Chinese and Japanese. (The cause is the training data — models learned from images carrying subtitles, watermarks, and signatures — not simply the presence of the word "dialogue" in your prompt.)
  • Layout is storytelling. Which panel is big, where the eye travels—those are authorial decisions. The model doesn't know your pacing, so it can't decide "this one goes full-bleed because it's the climax."

So the practical move isn't waiting for one-click generation. It's putting AI where AI is actually strong, which is where the five-stage split comes from.

The five stages, and who does what

StageWhat it decidesFit for AIWhy
1. Script / panel breakdownPanel count, what's in each panel, where dialogue landsYouIt's a storytelling decision, and it's the input to everything after
2. Composition roughPose, camera, where figures sit in frameVery strongCheap, fast to iterate, no style to worry about yet
3. Line artTurning the rough into clean linesStrongThe rough acts as an anchor, so results are stable
4. Finishing (screentone or color)Light, texture, moodStrongConsistency comes from shared references, not re-prompting
5. LetteringBalloon placement, type size, breathing roomYouText breaks, and balloon position is the pacing

⚠️ One terminology note, because it trips everyone up. What this article calls stage 1 is a written artifact — camera, action, setting, and dialogue as structured fields per panel. That's a script or panel breakdown. A storyboard (name / nemu in manga, thumbnails or layouts in Western comics) is the drawn version, with rough figures and panel borders, and it comes downstream of the script. AI write-ups routinely conflate the two, promising "AI storyboards" while actually describing prompt specs.

AI eats the drawing hours in the middle, not the judgment at either end. That's also why so many people report "I used AI and it wasn't faster"—they skip stage 1, let the model guess the story, and spend the savings sorting through images.

This is one division of labor, not an industry rule. Plenty of professionals do use language models at stage 1 to break prose into shot lists, and plenty find AI line art still needs hands and background perspective redrawn by a person. What follows is a route that works in practice, not the only one.

Stage 1: write a script the model can read

A storyboard is already the blueprint for a comic. What changes when the next reader is a model: it has to be written, and it has to separate the picture information from the story information.

Every panel needs at least these four fields. Leave one out and the model invents it.

  • Camera — close-up / bust / full body / wide, and low or high angle
  • Character and action — who's in frame, doing what, facing where
  • Setting — place, time of day, light source
  • Dialogue — in its own field. Never in the picture description, or the model will try to draw the words.

The bad version is "the protagonist stands sadly in the rain." No camera, no distance, no light source—so every generation returns something different. Rewrite it as "medium shot, slight high angle. The protagonist stands at the mouth of an empty alley. Night rain. The only light is the shop sign behind her." Consistency improves immediately.

Stage 2: lock the character before you generate panel one

This is the most important and most-skipped step in the whole pipeline: before generating any panel, give every character a fixed set of reference images.

In practice that means generating a clean character reference first — white background, no text, nothing else in frame — that pins down what this person looks like. From then on, every panel generation gets the same reference. Only then does the model have an anchor for "this is what this person looks like."

⚠️ How you feed the reference matters more than having one. Drop in a multi-view turnaround — front, side, and back on one canvas — and the common result is the model reading it as three people in the scene, or bleeding the back design onto the front. The reliable practice is one clean image, one view at a time; when you need more precision than that gives you, the next step is training a character-specific model rather than piling on more reference images.

Skip it and you land straight in the most common complaint: the protagonist in panel 3 and the protagonist in panel 17 look like siblings. Character consistency has a whole toolkit behind it—how many references, how to phrase descriptions, when it's worth training a dedicated model—covered separately. The point here is where it sits in the pipeline: it's setup, not damage control.

One more thing: keep appearance data separate from character writing. Height, build, hair color, eye color, usual outfit are for drawing. Personality, backstory, and motivation are for the story. Modern models do read natural language well, so a phrase like "wary, keeps her distance" can legitimately inform expression and posture — but prompt bloat is real, and concrete visual descriptors outrank abstract backstory when you need the same face twice. That split is the practical reason a character sheet has separate fields at all.

Stage 3: advance by stage; don't jump to a finished image

The beginner move is to ask for a full-color finished panel immediately and regenerate the whole thing whenever it's wrong. Two problems: it's expensive, and every regeneration also wipes the parts you liked.

The stable approach is to advance in stages, each one built on the last:

  1. Composition rough. Pose, body proportions, the big shapes of the background. No style, no detail. This is the cheapest stage and the one meant for volume—five or six tries for one panel costs you nothing worth protecting.
  2. Line art. Only once the composition is settled. Feed the character references here and the identity locks in.
  3. Finishing. Only once the line art is settled — and decide first whether you're making monochrome manga (screentone, blacks, effects) or color comics/webtoon (rendering). Those are different processes and different finished products; don't treat them as interchangeable. Style consistency comes from using the same references across the whole chapter, not from re-tuning a prompt every image.

⚠️ Be clear about the mechanism here. Diffusion models default to producing a fully rendered image in one pass; they will not politely hand you line art only. Getting the staged pipeline above requires explicitly constraining each pass — feeding the previous stage back in as the input image (image-to-image), or driving it with line-art control adapters. And there's an equally mainstream route running the other direction: you draw the composition rough, and AI does the line art and finishing. If you can already draw, that one gives you more control.

Keep every stage; don't overwrite. You will want to compare "was the line art better before rendering?", and when the same location comes back next chapter, the old composition rough is reusable as-is.

Stage 4: when it's wrong, decide refine vs. restart first

This is the single biggest time-saver, and most people only ever press "generate again."

What's wrongWhat to do
Composition, angle, and character are right; one detail is off (expression, hand position, one background object)Refine — instruct it to change only that and leave everything else untouched
The camera or pose is fundamentally not what you wantedRestart — go back to the composition stage; don't fight it on a finished image
The character's face doesn't look rightFix the input — not regeneration. A hundred more tries won't make it resemble them

"Face doesn't match, so keep regenerating" is the classic waste. Repetition doesn't teach the model who your character is. But fixing the input doesn't mean piling on more references either — past a point they dilute each other. It means swapping out a weak reference, raising the reference's weight, inpainting just the face, or training a character-specific model.

Stage 5: lettering stays human

There are two separate reasons here, and they get blurred together constantly:

  1. Technical. Have the model draw balloons and text into the image and you get broken or hallucinated glyphs — worse in Chinese and Japanese.
  2. Craft. Even with perfect text rendering, where a balloon sits determines where the reader's eye goes and how long the beat lasts. Balloon placement is part of page layout — artists compose negative space around intended balloon positions at the layout stage so key art doesn't get covered — while typesetting itself (font, leading, kerning, the lettering layer) is post-generation work in vector or lettering software.

The most robust practice is to overlay dialogue as its own layer on top of the art rather than letting the model draw words into the image. The benefits are concrete: edit a line without regenerating art, translate by swapping only that layer, and keep type size and line breaks adjustable. How to write dialogue that sounds like a person talking is its own subject.

Before / After: one-shot thinking vs. pipeline thinking

One-shot (❌)Pipeline (✅)
StartHand the model a story, ask for mangaWrite a structured script, four fields per panel
CharactersDescribe in each panel's promptLock one clean reference, feed the same one every panel
GenerationAsk for full color immediatelyComposition → line art → rendering
When wrongRegenerate everythingKnow when to refine, restart, or fix the input
DialogueLet the model draw the textOverlay as a layer—editable and translatable
ResultNice single images that never form a chapterA chapter that reads as one world

Three common mistakes

  • Generating before scripting. The model doesn't know your pacing, so it returns decent standalone illustrations. The written breakdown is the input, not an optional preamble.
  • Starting without a clean reference. This is the root cause of faces changing between panels, and fixing the reference costs far less than redrawing a chapter.
  • Only knowing how to regenerate. You erase the good parts and then spend the time you saved sorting images. Learning the refine/restart line beats switching tools.

The real fix: keep script, characters, and every stage in one place

The hard part of these five stages isn't any single step—it's that they have to be cross-checked constantly. Generating panel 17 means checking the character from panel 3; rendering means looking at the line art; editing a line means going back to the script. Spread that across separate tools and you spend the day switching windows and hunting for files.

LitMemo's manga image generation sits directly on top of the per-panel breakdown. Every panel has its own production pipeline—composition rough → line art → screentone → full color—and every stage's output is kept, so you can compare backwards. Each stage can be AI-generated or your own upload, and mixing the two is fine. The page shows the most advanced stage available, so while a chapter is in progress some panels can sit at full color and others still at line art — a production state, not how it ships.

Characters come in from character management: a character with no reference image is gated at the start and prompted to generate one, instead of letting you finish a chapter and only then notice the faces drift. Dialogue lives as a separate layer over the art, so edits and translations never require regenerating an image.

Summary

The right question about AI manga isn't "can it generate everything in one shot." It's "which of my stages can I hand off?"

The answer is usually the middle three: composition, line art, finishing. The two on the ends—script and lettering—are storytelling decisions, and giving them away tends to cost you time. And for the middle three to actually save hours, two things have to be in place first: a script written as four structured fields, and one clean character reference acting as an anchor. With those, the rest is advancing stage by stage and knowing when to refine instead of restart.

Frequently Asked Questions

Can AI generate a whole manga chapter automatically?
No, for three structural reasons: cross-panel character consistency depends on reference images you supply (the model doesn't remember the previous panel), text inside the image is unreliable (especially Chinese and Japanese), and panel sizing is a storytelling decision the model can't make because it doesn't know your pacing. The practical approach is the five-stage split, giving AI the composition, line art, and rendering stages.
Should I let AI draw the storyboard?
Not recommended — and separate the two artifacts first. The *storyboard* is the drawn version (thumbnails, panel borders, rough figures); what you hand a model is the **written script** upstream of it. That script is the input to every later stage—panel count, camera, and dialogue placement are all set there. It works best in four fields: camera (distance plus angle), character and action, setting (place, time, light source), and dialogue. Keep dialogue in its own field rather than in the picture description, or the model will try to draw the words.
Why does my character look different in every panel?
Because there's no fixed anchor. Each generation is an independent event, so text description alone can't reproduce the same face every time. The fix is to build a clean character reference first—white background, no text, nothing else in frame—and feed that same reference to every panel. Avoid dropping in a multi-view turnaround as-is: models frequently read front/side/back on one canvas as three separate people in the scene. It's setup, not a repair: regenerating a hundred times won't make a face start matching.
When a result is wrong, should I regenerate or edit?
It depends what's wrong. Composition and character right, one detail off (expression, hand position, a background object) → refine, instructing it to change only that and leave everything else alone. Camera or pose fundamentally wrong → go back to the composition stage rather than fighting it on a finished image. Face doesn't match → fix the input: swap out a weak reference, raise its weight, inpaint just the face, or train a character-specific model (not simply pile on more references, which dilute each other). The problem with full regeneration is that it erases what was already working.
Why generate in stages instead of going straight to full color?
Two reasons. Iteration cost: the composition stage only needs pose and background shapes with no style at all, so it's the cheapest place to try five or six versions of a panel; settling composition before line art and line art before finishing gives every stage an anchor and makes results more stable. Note this pipeline is not a model default — diffusion models return a fully rendered image in one pass, so each stage has to be constrained explicitly (feeding the previous stage back in as an input image, or driving it with line-art control adapters). And reversibility: keeping each stage lets you compare backwards and reuse an old composition when the same location returns next chapter.
Can I have AI draw the dialogue into the image?
The model will try, but the output is usually garbled or near-miss glyphs, and it's worse in Chinese and Japanese. The stronger practical reason is layering: keep dialogue as its own layer over the art and you can edit a line without regenerating the image, translate by swapping that layer alone, and adjust type size and line breaks afterwards. Balloon placement also controls the reader's eye and the length of the beat, which makes it a storytelling call in the first place.