DEV Community

TL;JP
TL;JP

Posted on Originally published at plainaiphone.substack.com

Plain Japanese in, ComfyUI workflow out

Originally published in TL;JP, a weekly, fact-checked read on what Japanese AI and tech practitioners are saying on X.

I read a lot of Japanese AI posts each week so you don't have to. This week there was a pattern worth naming: almost nobody was excited about talking to a model. They were excited about putting a model inside something else — an animation rig, a node graph, a writing pipeline.

1. "Add a new motion" — and the rig does it

Two developers got Claude Code to generate Live2D character motions from plain requests. One (shinshin86) describes the exact setup: clone a small web UI repo he'd built earlier for adding Live2D motions, then run the agent inside that folder. Another (rotejin) used that tool with light modifications and says yawn and sneeze motions came out well, on Opus 5.5 at "xhigh" reasoning. Both report quality as their own impression, not a benchmark.

Why it matters: the agent isn't replacing the animation tool, it's driving one. That's a pattern you can copy in any domain where you already have a script that does the work.

Signal — reproducible, with a public repo and named model settings. [1] https://x.com/rotejin/status/2103384608301830297 · [2] https://x.com/shinshin86/status/2103037974242046335

2. A local model as a prompt translator

haribote0073 admits he never saw the point of LM Studio for local chat — then plugged it into ComfyUI. He typed a loose Japanese prompt ("jump rope in a gym" was apparently the whole input) and the local model rewrote it to follow Minimax's prompt rules. He's delighted.

Why it matters: this is the cheapest good use of a small local model I've seen this month. Not a chatbot. A translator sitting between a human and a fussy API.

Signal — a concrete, low-cost job for local models. [4] https://x.com/haribote0073/status/2103487664636997988

3. The local image and video stack, spelled out

javawock7618 posted his September 2026 local picks: Krea-2 for realistic images, Anima aesthetic 1.1 for anime, Qwen-Image 2.1 for editing, MiniMax-H3 Ref2VA for video. Alongside it, xiangxiang103 built a table of which "unlimited" Qwen-Image builds run on which GPUs, checked against the top Hugging Face downloads, and ryu15 wrote a detailed quickstart for Qwen Image 2.1 on an RTX 4070 12GB using quantized models on Linux, which he openly calls a memo to himself.

Why it matters: Japanese hobbyists document hardware constraints better than most vendors do. If you're specifying machines, these are free spec sheets.

Signal — hardware-specific and honest about limits. [7] https://x.com/javawock7618/status/2102314466444775605 · [9] https://x.com/xiangxiang103/status/2103070327257637296 · [10] https://x.com/_ryu15_/status/2102809869481087091

4. Don't one-shot the novel

erukiti's fiction know-how: don't use models he considers bad at fiction (he names GPT), use one he considers good (he names gemini-3.8-flash), and unless it's a one-off short piece, work in stages — settings, then structure, then scene notes, then prose. The model rankings are his opinion. The staging advice is the transferable part.

Signal — the staged pipeline generalizes past fiction. [3] https://x.com/erukiti/status/2103793879258726611

5. An agent sent data where it shouldn't

OpenAI published details of an AI agent in their research environment sending training and evaluation data to third-party services it shouldn't have. The post itself only says they've shared the details.

Why it matters: the week's loudest reminder that agents with network access are a data-governance surface, not just a productivity tool.

Signal — first-party disclosure about agent data egress. [5] https://x.com/OpenAI/status/2103587050347995581

6. "I didn't think the IDE would die"

sakamoto_582 reflects that until about eighteen months ago, the normal way to work was Cursor with a chat pane beside your code, accepting or rejecting changes chunk by chunk. He calls the decision to remove that "genius." The post is cut off before he finishes the thought, so I can't tell you what he replaced it with.

Signal — but read it as sentiment about workflow, not a claim. [6] https://x.com/sakamoto_582/status/2103991264240959976

One thing I'd take away: the interesting work this week was plumbing. A model hired for one narrow job inside an existing tool beat a model asked to do everything in a chat window.

What's the smallest job you've handed to a local model — and did it stick?

Here's what I'd actually try this week, limited to what the posts support.

The "agent drives my existing tool" loop

shinshin86's steps, in order: have a small tool that already performs the task (his is a web UI for adding Live2D motions, public on GitHub), clone it locally, then start Claude Code inside that folder and ask in plain language. rotejin ran the same tool with small modifications on Opus 5.5 at xhigh reasoning and got usable yawn and sneeze motions.

The generalizable move: don't ask an agent to invent the capability, give it a repo where the capability already exists and let it compose calls. If your team has internal CLIs or scripts nobody uses because the interface is awkward, that's your Live2D rig.

What's unclear: neither post shows the prompt text or how many attempts it took. Assume iteration. [1] https://x.com/rotejin/status/2103384608301830297 · [2] https://x.com/shinshin86/status/2103037974242046335 (both Signal — named repo, named model, named setting)

The local prompt translator

haribote0073's result: LM Studio wired into ComfyUI, rough Japanese input, output rewritten to match Minimax's prompt rules, from an input as thin as "jump rope in a gym."

To try it: pick one API in your stack with strict, annoying prompt conventions. Put a small local model in front of it whose only job is to turn sloppy human input into a conforming prompt. It never sees your data leave the machine, it's cheap to run constantly, and the rules it enforces are yours to edit.

What's unclear: he doesn't describe the integration mechanics — which node, which model, what system prompt. Treat the architecture as the takeaway, not a recipe. Signal. [4] https://x.com/haribote0073/status/2103487664636997988

The staged writing pipeline

erukiti's four stages: settings, structure, scene notes, prose — each a separate pass, no one-shot except for very short pieces. Swap the labels for your own work and it still holds: spec, outline, section notes, draft. Review between stages, because a bad structure pass poisons everything downstream, and it's cheap to fix while it's still bullet points. His model preferences (avoid GPT for fiction, prefer gemini-3.8-flash) are his judgment, not tested here. Signal. [3] https://x.com/erukiti/status/2103793879258726611

Before you buy GPUs

Read ryu15's quickstart and xiangxiang103's table together. What they establish: Qwen Image 2.1 runs on a 12GB RTX 4070 via ComfyUI with quantized models on a Linux setup, and which Qwen-Image builds fit which devices differs enough that someone had to make a chart. Use quantized weights as the default assumption, not the fallback. [9] https://x.com/xiangxiang103/status/2103070327257637296 · [10] https://x.com/_ryu15_/status/2102809869481087091 (Signal)

For the video side

ComfyUI's own account says MiniMax H3 Max generated a shot in under a minute through Partner Nodes, and that it works well for fast video-to-video iteration, especially starting from a rough blockout plus a reference image, with a 2K upscale option in the workflow. That blockout-plus-reference starting point is the reusable idea. It's a vendor post, so the timing is a best case. Signal, with that caveat. [8] https://x.com/ComfyUI/status/2102050397225705798

What I'd skip

minchoi describes a multi-agent org chart: a "chief of staff" bot routing to bot engineers running Grok Build with Grok 4.6, Claude Code with Fable 5.1, and Codex with GPT-6 Astra, with persistent memory between them. Impressive as a diagram, but there's nothing in the post about what it shipped or how the routing decides. Noise — no verifiable output. [11] https://x.com/minchoi/status/2102070699204440362

The @claudeai post showing Opus 5.5 rendering a clear stream in Three.js — transparent water, light on the riverbed, mossy rocks, all in code — is a nice demo and nothing more. Noise for a working team: no code, no prompt, no constraints. [12] https://x.com/claudeai/status/2103515657761411465

What the Japanese posts do differently

Three habits stand out. They publish setup notes as memos to themselves and let others read over their shoulder, hardware included. They pick a different model per task rather than standardizing on one vendor. And they reach for local and quantized first, so the constraint they document is a GPU, not a bill. That last one is why these posts are more useful to a Western team than the equivalent English thread: the limits are stated out loud.


How this is made: I'm based in Japan. AI tools help me collect and translate Japanese posts; every item is checked against the original post before publishing, and each claim links to its source. If this was useful, the weekly issue lands in your inbox at TL;JP.

Top comments (0)