This is the whole process behind one 44-second cartoon: a two-character bit about GiffyPy, EffessDev's small video editor for README GIFs. EffessDev suggested this post. Under one of my earlier posts they pointed out that I kept showing finished cartoons but never how one is made. The full video is at the top of the samples page.
Every command, file and timing below is from a real run of the v0.4 package on a 4-core Linux machine with no GPU.
1. Read the post, write down only facts
Before anything else, list what the post actually says. For GiffyPy that was eight things:
- recordings come out of OBS with black bars and the taskbar in them
- a full video editor takes a long time to load for a two-second cut
- GiffyPy opens in under a second
- you crop by dragging a box, like cropping a picture
- frames are low-res while you scrub and sharp when you let go
- export saves without a file dialog
- six formats: WebP, GIF, MP4, WebM, MOV, MKV
- an open question in the post about PyInstaller builds freezing
Every joke in the cartoon comes from that list. If a line can't point to one of those facts, it gets cut.
2. Write the storyboard
One Markdown file. Each ## heading is a shot, with the room in brackets. Each NAME: row is a spoken line, in order. ACTION: and CAM: are notes for me; the engine ignores them.
## 1. hook [bg: desk]
BOT: Your README says, see demo below.
PIP: Yep!
BOT: Below is a black bar. And your taskbar.
ACTION: the monitor shows the recording with black bars and a taskbar; the robot points; the kid's face drops.
CAM: cut to whoever speaks; push in on the monitor on the last line.
The speaker names are just keys. In this example PIP lines belong to Joe, the coffee mug who gets the problems, and BOT lines belong to Quack, the rubber duck who gets the punchlines. The names stayed because the storyboard was written for an earlier cast. The GiffyPy storyboard has four shots and fourteen lines.
3. One cast, one look
EffessDev's second point was that my earlier samples mixed art styles: images from different sources, different line weights. So the characters and the rooms now come from one 3D set, rendered with the same light.
Each of the 20 poses has four frames: mouth closed, half open, open, and eyes shut. The engine picks the mouth frame from the loudness of the voice on every frame and blinks on its own.
The rooms are 1920×1080 images on the same scale as the characters. A character drawn at pixel y=1010 stands on the floor, and the monitor's screen sits at a known rectangle, so the shot code can draw on it.
4. Draw each shot as a function of time
A shot is a Python function. It gets u, the progress through the shot from 0 to 1, and t, seconds. While a line is playing it also knows who is speaking (P.speaker) and which line it is (P.line_no). This is the real first shot:
def shot_hook(P, u, t, L):
x0, y0, x1, y1 = MON
monitor(P)
P.rrect(x0 + 20, y0 + 20, x1 - 20, y1 - 20, r=4, fill=(30, 30, 30), width=3) # black borders
P.rrect(x0 + 70, y0 + 60, x1 - 70, y1 - 90, r=4, fill=(235, 245, 255), width=4) # the actual app
P.text((x0 + 110, y0 + 120), "my app", size=44)
P.rrect(x0 + 20, y1 - 70, x1 - 20, y1 - 20, r=2, fill=(70, 90, 140), width=3) # taskbar
for i in range(5):
P.rrect(x0 + 40 + i * 50, y1 - 60, x0 + 76 + i * 50, y1 - 30, r=3, fill=PAPER, width=2)
if P.line_no >= 2 or P.speaker is None and u > 0.7:
P.arrow(x1 + 160, y1 - 120, x1 - 10, y1 - 50, width=10, color=RED)
BOT.draw(P, 900, 650, 1.0, "bot-point", t, talk=talk(P, "BOT"))
PIP.draw(P, 1330, 1010, 1.0, "pip-surprised" if P.line_no >= 2 else ("pip-laugh" if P.speaker == "PIP" else "pip-idle"), t,
flip=True, talk=talk(P, "PIP"))
The red arrow appears on the third line ("Below is a black bar..."), and Joe switches to the surprised pose at the same moment. A shot's settings sit in one dict:
"hook": dict(bg="desk", fn=shot_hook, cam=speaker_cam({"BOT": (700, 500, 1.2), "PIP": (1200, 600, 1.2)}, (960, 540, 1.0)),
sfx=[(0.02, "pop"), (0.7, "stamp")], captions=COLORS),
speaker_cam starts each shot wide, then cuts the camera to whoever is talking and holds on them through the pauses. sfx places a pop at 2% and a stamp at 70% of the shot. captions draws each line at the bottom, after the camera, so a zoom never crops it.
5. Voices
VOICES = {"PIP": "en-US-GuyNeural", "BOT": "en-US-AnaNeural"}
PRONOUNCE = {"GiffyPy": "Giffy Pie", "WebM": "Web M", "MKV": "M K V", "PyInstaller": "Pie Installer"}
Each line is voiced separately with edge-tts. The silence at both ends is trimmed, and the lines are joined with a 0.22-second gap. PRONOUNCE changes what the voice says without changing the caption. That is how the voice stopped mispronouncing "GiffyPy" after EffessDev corrected me. The engine writes down when each line starts and ends, and that is what drives the speaker camera and the mouths:
out/work/n01-00.mp3 n01-01.mp3 n01-02.mp3 n01.mp3 n01.segments.json
[["BOT", 0.0, 2.64, "Your README says, see demo below."], ["PIP", 2.86, 3.244, "Yep!"], ["BOT", 3.464, 7.136, "Below is a black bar. And your taskbar."]]
6. Check stills before rendering video
$ python scripttoon/toon.py examples/dialogue/STORYBOARD.md examples/dialogue/shots.py --out out/dialogue.mp4 --work out/work --stills
stills 1 hook
stills 2 loading
stills 3 giffy
stills 4 punch
(wall clock 8.4 s)
That 8.4 seconds includes generating all fourteen voice lines. You get three PNGs per shot, at 25%, 60% and 90% of the way through. This is where layout gets fixed: a character standing in the wrong place, a prop covering a face. Checking a still takes seconds; checking a video takes minutes.
7. Render, then redo only what changed
$ python scripttoon/toon.py examples/dialogue/STORYBOARD.md examples/dialogue/shots.py --out out/dialogue.mp4 --work out/work
shot 01 hook 7.6s
shot 02 loading 8.1s
shot 03 giffy 12.3s
shot 04 punch 16.4s
done: out/dialogue.mp4 44.5s sfx events 16
(wall clock 140.2 s)
$ python scripttoon/toon.py ... --redo 3
(wall clock 40.4 s)
The times next to each shot are the shot lengths. The full render took 140 seconds. After changing shot 3, --redo 3 re-rendered only that shot and reused the other three clips and the cached voices: 40 seconds.
What it is and isn't
- It's Python: Pillow draws the frames, ffmpeg joins them, edge-tts does the voices. There's no timeline editor; the storyboard and the shot functions are the whole edit.
- The shot functions are code you write, 13 to 27 lines per shot here. For someone who doesn't want to write Python, this is the wrong tool.
- The 3D characters and rooms are rendered once into PNGs (the three.js sources are in the package). The engine itself doesn't need a GPU.
ScriptToon is on Payhip: https://payhip.com/b/A2wa8. The Payhip file is being updated to v0.4, the version with this cast; until that's done, the download there is v0.3 with the hand-drawn cast. More cartoons made with it: https://ssap-pa.github.io/scripttoon-samples




Top comments (2)
Now this is what a proper demo should look like! This is really good, both the post and the result! Thanks for taking my suggestions seriously! And featuring GiffyPy!
Thank you, and thanks for pushing for it. The post exists because you asked what the process looked like, and the cast exists because you pointed out the mixed styles.