Where generative AI video breaks: field notes from a Taipei studio
Most of the discussion about AI animation over the past year has circled one question: will it replace animators.
I don't have an answer to that, and I don't think anyone can answer it yet.
The more useful question is where it breaks.
I started UNI Entertainment in Taipei in 2023. We put generative tools into real jobs with real delivery dates, not into test films. A year of that leaves you with some unromantic notes. These four we found the hard way.
Some styles the model cannot hold
Image-to-video models struggle with flat vector art.
Silhouettes survive. A character outline moves more or less correctly. But anything that has to hold its shape starts to fall apart: geometric blocks, regular grids, lines that need a consistent weight. Edges wobble. Grids skew. The proportion of a colour block drifts on its own.
The reason is not mysterious. These models predict the next frame at the pixel level. They do not know that the blue area is a shape, and nothing in the process gives them a reason to. Your eye does that part on its own.
So we stopped forcing it. For shots in that style we animate programmatically, and a straight line stays straight. We save the generative tools for what they are actually good at, which is depth, material and light.
The defaults will bite you
This one is boring, and it still made us redo work.
Platform defaults are not always what you want. In the version we use, video length defaults to 5 seconds when the real range is 3 to 15. Quality defaults to standard 720p. Sound defaults to on, so the model scores your clip by itself, and that audio ships with the file if nobody checks. Image models are the same story: resolution defaults to 1k when 2k and 4k are available.
Small things. But you find out a batch of sixty shots came back at 720p, and that is a day.
The defaults also change between versions. So our rule now is that every parameter gets written out explicitly, none of them left to default, and we re-confirm them at the start of each job.
Revision does not work the way it used to
This is the one that changed how we work the most.
With an animator you can give a local note. Change the third character's shirt to red, leave everything else alone. They understand you, and they can do it.
A generative model does not work that way. Ask it to change three objects to four and it re-rolls the whole image. What comes back is not the same picture with one more object in it. It is a different picture. The composition moved, the light moved, and the angle you spent the last round getting right is gone.
So the way you give notes has to change. Two things we learned.
Batch the changes. If it can be said in one round, don't spread it over three, because every round is another roll of the dice.
And when you talk about colour, say what it cannot be. Ask for "a different colour" and the model circles the same few you just rejected. Say "not orange, not ochre" and it finally moves.
That sounds like learning how to talk to a tool. It mostly is. It has also become a skill worth practising.
Chinese speech recognition is not a record
We work on Chinese-language jobs, so we use speech recognition to write up recordings and dictation. There is a trap here.
These models process in chunks of roughly 30 seconds. On Chinese, words go missing at the chunk boundaries, and nothing tells you they went. The transcript reads perfectly well. It is just missing a sentence.
The default output is also Simplified, so it needs converting to Traditional.
Our rule now is that a transcript is a draft, not a record. If it has to serve as the record, a person listens to it again.
Back to the first question
I have said many times that AI accelerates creativity, and I still think so.
But "accelerate" has a shape once you put it inside a real pipeline. It accelerates execution, not judgement. The cost of producing early work has dropped. The cost of choosing has not, and it has gone up, because now there is more to choose from.
What is scarce has moved. It used to be whether you could make the thing at all. Now it is whether you can tell which one is right.
Tools change. Taste doesn't.
