← Back to Tutorials Creative & Design

Assets First, Then Animate

Everyone's copying the Seedance styles. The creators actually landing hits are copying something else first. Four working prompts, and the one step nobody names.

Creative & Design⏱ 12 min read● Intermediate

What's Actually Trending

If you've been anywhere near AI video in the last few months, you've seen the output. Manga panels turning into animated anime. A single photo becoming a fantasy warrior. Eight-scene short films where every scene is a completely different animation style. Somebody's cat, rendered as claymation, screaming.

Most of it is coming from one model: ByteDance's Seedance 2.0.

Here's the thing worth knowing before you spend a dollar on it. Seedance 2.0 is not hype. On Artificial Analysis, which runs blind head-to-head human preference tests, Seedance 2.0 is currently the highest-rated image-to-video model in the world. It beats Google's Veo 3.1 and Kling 3.0 on both the image-to-video and text-to-video boards, at roughly a third to half their price per minute. That's an independent leaderboard, not a press release.

So the model is real. But here's where most people go wrong, and it's the reason this article exists.

Everybody looks at the viral output and asks "what was the prompt?" They copy the style words. They paste in "Studio Ghibli style, cinematic, breathtaking, 8K" and they get slop, and they conclude the model is overhyped or that they're bad at this.

The prompt was never the trick.

The real pattern: Go read the actual production breakdowns from the teams landing these hits. Every single one does the same thing before writing a video prompt. They build reference assets first. Nobody calls it a trend, because it isn't a style. It's a workflow.

That's what this article teaches. We'll cover what's actually spreading, how to get access if you're in the US (there's a wrinkle there), the asset-first workflow, why your prompt is probably built wrong, and where the legal line sits. By the end you'll have four working prompts and, more usefully, the reason they work.

The Thing Every Viral Seedance Piece Has in Common

Let's look at two real productions.

The eight-style short film

Higgsfield's team made a short film where a man teleports through eight worlds, each in a different animation style: anime, Moebius-style French graphic novel, chibi low-poly, cyberpunk claymation, 1980s black-and-white seinen manga, clay hell, a 2D cartoon dropped into a live-action wedding, and a live-action ending.

Read their breakdown and the first instruction has nothing to do with animation:

From the production notes"Every good film starts with consistent assets. Before generating any scenes, create a character keyframe and a watch prop sheet. The watch appears in every single scene, so getting it right first saves you a lot of time."

A prop sheet. For a wristwatch. Before a single frame of video. That's the whole lesson in one line, and most people skip straight past it to the fun part.

The K-pop sci-fi series

Same team, bigger swing: a fully AI-generated sci-fi trailer with five original characters, mech suits, a monster invasion, and a K-pop dance sequence. Roughly a hundred generations made the final cut. Hundreds more didn't.

Their character pipeline ran three steps, per character, before any video existed:

Step 1: The face Generate the face alone, on a white studio background. Run it repeatedly until the bone structure and eyes are right. Nothing else in frame.
Step 2: The outfit Generate the outfit separately. Every detail that matters gets written explicitly here, before it becomes a problem across 40 shots.
Step 3: Fuse them Merge face and outfit into one master character sheet. This is the asset Seedance checks against every time that character appears.
Then, and only then Write the video prompt, referencing the finished asset. Five characters meant five pipelines. Cast locked before camera rolls.

Their words: "Run this pipeline once per character. Five master assets. Cast is locked."

Three-panel character reference pipeline: an isolated neutral headshot on white, the outfit alone as a flat lay, and the two fused into a full-body master reference sheet.

The part that should convince you

Here's what makes this more than one company's house style. China's micro-drama industry went almost fully AI in about ninety days. In the first quarter of 2026, roughly 128,000 micro-dramas were released nationwide and over 95% of them were AI-made. A production that cost $200,000 now costs $7,000 to $14,000.

And a brand new job appeared on Chinese job boards, hundreds of listings deep, most requiring no prior industry experience. The title translates roughly to AI asset curator. The job: read the script, and generate the reference images of characters, costumes, and scenes for the video model to anchor to.

An entire industry, working independently, at enormous scale, invented a full-time role for the step everyone else skips. That's not a coincidence. That's the actual bottleneck revealing itself.

Internalize this one: Seedance is not a slot machine you feed words into. It's a renderer that needs something to render against. The prompt describes the motion. The assets decide what things look like. If your character drifts between shots, that is almost never a prompt problem.

Getting In

There's a wrinkle here, and it's poorly covered elsewhere, so let's be direct about it.

BytePlus, ByteDance's own enterprise platform, is not available in the United States. That's not a slow rollout or a waitlist. Their own terms say the services "are not available in the United States," and the US is absent from their published availability list of roughly 170 countries. Canada, the UK, the EU, Japan, and Australia are all on it. The US isn't.

So if you're American and you want to use Seedance, you go through a platform that has its own arrangement with ByteDance. You're a customer of that platform, not of ByteDance. Higgsfield is a San Francisco company with an official ByteDance partnership, which is why it's the route we're using here.

Prices as published by Higgsfield in June 2026, comparing a standard 720p 8-second clip across the main options:

Platform Per clip Subscription Catch
Dreamina (CapCut) $1.29 $18 to $82/mo Regional gaps, daily token ceiling
Higgsfield $1.55 $9 to $129/mo Web only, no API
Magnific $1.86 $7 to $249/mo Design tool first, video second
fal.ai $2.43 Pay per use API only, no interface
Runway $2.88 $15 to $95/mo Priciest, updates arrive late

Two honest notes on that table. First, it's published by Higgsfield, who obviously have a preferred answer, so treat the framing with the skepticism you'd apply to any vendor comparison. The numbers themselves are checkable and they held up when we spot-checked fal.ai's public rate of $0.3034 per second.

Second, Dreamina is genuinely cheaper per clip and gets model updates first, being ByteDance's own. If you're outside the US and outside the affected European regions, it's a reasonable pick. The daily token ceiling is the thing that bites at real volume.

Why Higgsfield for this workflow specifically: the asset-first pipeline needs an image generator and a video generator in the same place. Higgsfield bundles Nano Banana Pro, Soul Cinema, and Seedance under one credit pool, so your character sheets and your video draw from the same balance. Doing this across three tools with three logins is where the workflow falls apart in practice.

The Higgsfield Seedance 2.0 generation panel, showing the reference upload slots for image, video and audio, the prompt field, and the resolution and duration controls.

Budget reality: the $129/mo Ultra plan gives about 83 standard clips or 107 fast clips a month. That sounds like a lot. It isn't. The K-pop trailer burned roughly a hundred generations for a piece measured in seconds, and hundreds more that didn't make it. Plan for a hit rate, not a headcount.

Build the Cast Before You Roll Camera

Now the workflow itself. Three stages.

Stage 1: Generate your assets in isolation

One thing per image. Face on a white background. Outfit on a white background. Prop on a white background, ideally as a multi-angle sheet. Location as its own frame.

ByteDance's own documentation is unusually blunt about why this matters. Two named causes of character drift, in their words:

From ByteDance's Seedance 2.0 prompt guide"Mixed reference images: Providing the model with a single combined image that includes the face reference, full-body/half-body pose reference, outfit reference, detail reference... Face ratio is too small: In mixed reference images, the face area accounts for too small a proportion... the model assigns insufficient weight to them."

Their prescribed fix, again in their words: prepare "a close-up face image (headshot, retaining only the face; no expression is best; minimize interfering elements such as shoulders, neck, and background)."

No expression is best. That's counterintuitive and worth sitting with. You want a neutral face, because you're giving the model an identity to work from, not a performance to copy.

Stage 2: Fuse into a master sheet

Combine your isolated assets into one production reference. Higgsfield's fusion prompt is almost comically plain:

"The character from image 1 wearing this outfit from image 2. Full body shot, white studio background, soft cinematic lighting, realistic."

That's it. Fifteen words. The complexity lives in the assets, not the sentence.

The mistake almost everyone makes: ByteDance says outright that "using multi-view character images is not recommended." Multi-view sheets contain the same character at different angles, and the model may read those angles as different people, which makes identity drift worse. Your instinct to hand it a turnaround sheet is wrong. One clean headshot beats six angles.

Stage 3: Write the video prompt against your assets

Only now do you write motion. And you write less than you think.

Seedance 2.0 accepts up to 9 images, 3 video clips, and 3 audio files at once. Do not use all of it. ByteDance's own recommendation:

"Recommended configuration (4 to 5 assets in total): 1-2 character images (facial close-up / full body) + 1 scene image + 1 camera movement video + 1 audio clip. It is not recommended to use the full asset limit. Too many assets will make it difficult for the model to judge feature priorities."

Two more rules from the same doc that carry real weight. Place important assets first, because the earlier an asset appears in the prompt, the more precisely it gets referenced. And keep reference people at four or fewer, because past four, output stability drops.

The shot-planning rule worth stealing: the K-pop team introduced every character exactly twice. Once establishing who she is, once establishing what she can do. That's it. A rule that simple is what keeps a hundred generations from becoming incoherent.

Why Your Prompt Is Probably Wrong

Here's a prompt of the kind that circulates in every AI prompt pack on social media. It's for a Ghibli-style scene:

What most people write"The video opens with a breathtaking forest scene, where the camera pans over a lush, vibrant landscape bathed in soft, warm sunlight. Towering trees stretch toward the sky, their leaves rustling gently in the breeze. A young adventurer, with a joyful expression and wild, untamed hair, walks along a moss-covered path... The camera follows them as they move deeper into the woods, passing by tranquil streams and ancient rocks... The scene feels timeless, evoking a sense of wonder and innocence. As the adventurer stops to gaze at a magical, glowing tree, the camera zooms in on her face..."

It reads beautifully. It will not produce what it describes. Six reasons.

It's a treatment, not a prompt. "The video opens with" means you're narrating a finished film to a person. Seedance isn't an audience. ByteDance says it plainly: a good prompt "is not simply 'copywriting-style description', but an engineering-style instruction."

Three camera moves in one shot. It pans, then follows, then zooms. The official guidance: "Try to specify only 1 type of camera movement in a single shot. Do not require push, pull, pan, and move at the same time, as this will increase image instability."

It's wildly over budget. Count the beats: landscape pan, walking, creature appears, moving deeper, streams, ancient rocks, stop at tree, zoom on face. That's about forty seconds of story. Seedance 2.0 caps at fifteen. Something gets rushed or dropped, and you don't get to choose which.

The emotion is unrenderable. "A sense of wonder and innocence" is not a thing a camera can see. ByteDance ships an actual lookup table for this, and it's the most useful thing in their docs:

Don't write Write this instead
Sadness lowering the head, shoulders trembling slightly, eyes reddening, fingers unconsciously clutching the corner of clothing, tears welling in the eyes but not falling
Nervousness frequently checking the watch, fingers constantly tapping the tabletop, rapid breathing, eyes darting away, unconsciously biting fingernails
Anger both fists clenched, jawline tense, chest heaving violently, eyes as sharp as knives, squeezing words out through gritted teeth

The audio is wasted. "The soundtrack is filled with delicate piano notes" is prose. Seedance 2.0 has literal special-character syntax for this, and most users have no idea it exists:

The syntax nobody uses( ) music → (fast-paced rock music is playing in the background) < > sound effect → < dog barking can be heard in the distance > { } dialogue → {Hello, world} 【 】 subtitles → 【Chapter One: Departure】

There are no constraints. This is a named, required element of the official formula, and it's missing from essentially every prompt pack in circulation. ByteDance: "Constraint words are very important. They can effectively avoid visual flaws, deformities, breakdowns, and unreasonable elements."

Here's the official formula, verbatim:

The actual formulaprecise subject + action details + scene/environment + lighting & color tone + camera movement + visual style + image quality + constraints

Four Styles, Four Prompts

These are built to the formula above. Three shots each, one camera move per shot, audio in native syntax, constraints block, and they fit inside fifteen seconds. Paste the block, not the heading.

1980s anime

1980s cel anime, bold black linework, cross-hatching, screentone dots, high-contrast shadows, warm pink-orange sunset grade, visible film grain. Shot 1: Wide shot. City skyline at dusk, pink and orange sky. A woman in a long coat stands at the edge of a rooftop, back to camera, coat hem moving in the wind. Camera holds static. Shot 2: Medium shot, three-quarter angle from behind. She checks a folded paper map in both hands, then lowers it. One hand pushes hair back from her forehead. Camera pushes in slowly. Shot 3: Low angle from the street. She descends a fire escape, two steps at a time, and steps onto the pavement. A neon sign flickers behind her. Camera holds static. (soft synthwave, slow pulse) < wind, distant traffic, footsteps on metal > Constraints: no subtitles, no text, no watermark. Single character only, no duplicate figures. Face remains stable without deformation. One camera movement per shot. Natural motion, no stutter or flicker.
Output frame from the 1980s anime prompt: a pink and orange sunset sky over a city skyline, rendered in bold cel-anime linework.

Do not ask for "choppy frames." Good instinct, since real 1980s anime is animated on twos and genuinely is choppy. But Seedance runs locked at 24fps and can't animate on twos. It reads "choppy" as stutter, which is a documented failure mode, not a style. You'll get artifacting and blame the model. The linework and screentone carry the era instead.

An earlier version of this prompt said "a young hero with spiky hair" who leaps from the roof and lands in a crouch. It named no studio, no character, and no celebrity. Higgsfield rendered it and then blurred it behind a rights gate: "This output may contain copyright or likeness-protected content." Spiky-haired figure plus rooftop plus jacket plus neon at dusk is a silhouette the filter recognizes, whatever you call it. The version above trades the hair and the leap for a coat and a fire escape. Section 08 covers why that works.

Illustrative 2D

Children's picture book illustration, soft rounded outlines, flat vibrant color fills, gentle paper grain, storybook proportions, warm daylight. Shot 1: Wide shot. A small hillside town of colorful simplified buildings. An orange tabby cat trots along a cobblestone street toward camera, tail up. A few leaves drift past. Camera holds static. Shot 2: Medium shot, side-on, tracking with the cat as it walks. It glances up at a baker setting bread in a window. Trees sway gently behind. Camera tracks laterally at the cat's pace. Shot 3: Medium shot. The cat stops at a flower stall, sits, and paws once at a low hanging bloom. Its ears flick forward. Camera pushes in slowly and holds. (light playful acoustic, simple melody) < soft breeze, distant birdsong, faint footsteps on stone > Constraints: no subtitles, no text, no watermark. Single animal subject, no duplicate figures. The cat remains the sole subject throughout. One camera movement per shot. Natural motion, no flicker.
Output frame from the Illustrative 2D prompt: a hillside town of colourful simplified buildings with terracotta roofs, drawn in a children's picture book style.

Two things got fixed here, and the second one cost us a generation to learn.

The first version handed the camera off to the flowers halfway through, so they glowed and danced while the girl was still meant to be the subject. Two subjects competing for focus is a known breaker. Seedance wants one primary action with supporting environmental motion around it.

The second version fixed that and still failed, flagged as sensitive content before it rendered a frame. The subject was a young girl, and shot 3 put a medium close-up on her with "her eyes close, shoulders drop slightly." A minor, in close-up, with a described physical reaction is the exact shape an automated classifier is built to catch. It was over-triggering. It was also, given what it is built to prevent, over-triggering correctly.

Swapping the girl for a cat isn't a dodge, it's a better prompt. Animals are native to picture books, they carry the storybook read more strongly than a human does, and ears and paws externalize emotion just as well as eyes and shoulders.

The technique and the filter are in tension. Section 05 tells you to replace abstract emotion with physical tells, and that advice is right. But when the tells are lips, eyes, and shoulders, and the subject is a child or a young woman in close-up, you have written something a classifier will read badly. Use adults, use animals, or keep the tell gestural. Hands in pockets. Ears forward. A hand that rises halfway and stills.

Soft watercolor forest

Hand-painted watercolor backgrounds, soft cel shading, pastel palette, visible brush texture, dappled sunlight, dreamy soft edges. Shot 1: Wide shot. A sunlit forest clearing, tall trees, moss-covered path. A young woman with untamed hair walks slowly into frame, boots pressing into soft earth. Leaves move gently overhead. Camera holds static. Shot 2: Medium shot from behind, following her at walking pace as she moves deeper into the trees. Small motes of light drift up from the moss around her feet and hang in the air. Camera follows steadily. Shot 3: Medium shot, low angle, framed from behind her shoulder. She stops and tilts her head back to look up. One hand rises halfway, then stills. Warm light moves across the canopy above. Camera pushes in slowly and holds. (delicate solo piano, sparse notes) < leaves rustling, distant birdsong, quiet footsteps on moss > Constraints: no subtitles, no text, no watermark. Single character only, no duplicate figures. Face remains stable without deformation. One camera movement per shot. Natural motion, no flicker.
Output frame from the soft watercolor forest prompt: a sunlit forest clearing with tall trees and dappled light on a moss-covered path, a lone figure small in frame.

Shot 3 originally read "medium close-up on her face... her lips part slightly, eyes widen." Same trap as the prompt above, so it's now an over-the-shoulder with a stilled hand. It reads as wonder without pointing a camera at a young woman's mouth.

You'll also notice this one isn't called the Ghibli prompt. Section 08 explains why, and it's the most useful thing in this article if you plan to sell anything you make.

Polished 3D street scene

Polished 3D animation, soft global illumination, warm bounce light, subsurface scattering on skin, detailed cloth and hair simulation, shallow depth of field, saturated daylight grade. Shot 1: Wide shot. A sunlit city street decorated with hanging bunting and paper lanterns, a few people moving in the soft-focus background. A man in a bright patterned jacket walks toward camera, looking around at the decorations. Camera holds static. Shot 2: Medium shot, tracking with him at walking pace as he weaves between two market stalls. Fabric and hair move naturally with his motion. Camera tracks laterally. Shot 3: Medium shot, low angle. He stops in front of a street musician playing a violin. His eyebrows lift, he grins, and he slides his hands into his pockets. He rocks once on his heels. Camera pushes in slowly and holds. (upbeat live violin, warm and lively) < street ambience, soft chatter, violin playing close > Constraints: no subtitles, no text, no watermark. Single main character, no duplicate figures. Face remains stable without deformation. Background figures stay soft and out of focus. One camera movement per shot. Natural motion, no flicker.
Output frame from the polished 3D street scene prompt: a sunlit street hung with paper lanterns and bunting, market stalls either side, a figure walking toward camera.

The first version had a child running through a crowd to watch a street performer juggling. All three details had to go. Sprinting is explicitly on ByteDance's avoid list. Juggling is multiple objects in ballistic arcs with hand contact, in a model whose own docs admit weakness in "physical stability in complex motion," which made it the single most likely thing across four prompts to visibly break. And the child in close-up was heading for the same sensitive-content flag that killed the picture book prompt. An adult, a violinist, and hands in pockets keep the beat and clear all three.

What Breaks, and How to Fix It

ByteDance publishes a candid failure FAQ, which is rare and worth reading. Three admissions in it stand out because they say a problem "cannot be avoided 100%." Here are the fixes that matter most.

Problem Fix
Subtitles you never asked for Can't be fully prevented. But: "prioritize generating videos in landscape (the probability of generating subtitles in landscape is significantly lower than in portrait). You can later crop it to portrait."
Duplicate characters, the "twin" effect Bind each role to one image. Use single-person references. And: "Do not directly use the complete script as the prompt. Overly redundant copy can easily cause confusion."
Jump cuts when extending a clip Post-production, not prompting. Trim 6 frames off the end of the previous segment and 1 frame off the start of the next. Repeat at every seam.
Quality decay on repeated extension "Multiple continuations will compound the degradation, and mottled color blocks are especially likely to appear in character face regions." Limit your continuations.
Wrong number of people Keep reference people at four or fewer per image, then image-to-video from there.
Plasticky skin Add "no 3D, no cartoon, no VFX" to force realism.
Bland, generic output Cut adjective stacking. "Beautiful, stunning, gorgeous light" becomes one strong word. Clarity beats intensity.

And one that fails silently, which makes it the nastiest: conflicting instructions produce garbage rather than an error. The model cannot reconcile "peaceful meditation garden with loud rock concert and quiet library atmosphere." It won't tell you. It'll just hand you something wrong and take your credits.

How to iterate without burning money: prototype at 480p, ship at 720p. Lock your seed, because without seed control you cannot tell whether a change came from your edit or from randomness. Change one phrase at a time. And expect two or three takes as normal. Treat it like casting, not a one-shot.

The Line Between Style and Theft

You need the background here, because it explains the filters you're going to run into.

Seedance 2.0 launched on 12 February 2026. Within 24 hours, an AI clip of Tom Cruise fighting Brad Pitt on a rooftop had millions of views. Within four days, Disney, the Motion Picture Association, Paramount Skydance, Warner Bros., Netflix, and Sony had all sent cease-and-desist letters. The MPA's was its first ever to a major AI company. Disney called Seedance "a pirated library of Disney's copyrighted characters." SAG-AFTRA condemned it. By March, two US senators were demanding ByteDance shut it down, and the global rollout was paused.

ByteDance responded with C2PA content credentials, visible and invisible watermarks, face detection filters, and copyrighted-character blocking. Those filters are in the model. They don't care which platform you're on.

Which brings us to the useful part.

The filter reads what's in the frame, not what you called it. This is the insight most guides miss entirely. People obsess over whether they can type "Ghibli" or "Pixar." That's the wrong worry.

Take the original Ghibli prompt from Section 05. It never actually says the word "Ghibli" in the body. It describes soft watercolor and pastel light. Sounds safe.

But look at what it puts on screen: a young girl, in a forest, when a small glowing creature appears beside her, and they walk to a magic tree. That's a Totoro composite. The style label was never the risk. The ingredients were.

Our rewrite swapped the creature for drifting light motes. Same feeling of wonder, no character to recognize.

This happened to us while writing this article

We ran the four prompts from Section 06 through Higgsfield before publishing. The 1980s anime one, in its first form, read: "a young hero with spiky hair stands at the edge of a rooftop... steps off the roof edge and drops through frame. A neon sign flickers in the background."

No studio. No character name. No celebrity. Written deliberately clean, by someone who had just spent a week researching exactly which words trip the filters.

Seedance rendered it. Then Higgsfield blurred it and put a gate over it: "Rights verification required. This output may contain copyright or likeness-protected content." Underneath, a button reading "I own rights to this content."

Nothing in the prompt named anything. But a spiky-haired figure on a rooftop in a jacket, against a neon skyline at dusk, in 1980s cel style, is a silhouette. The filter isn't reading the sentence. It's looking at the picture the sentence made.

That button deserves a hard look. "I own rights to this content" is not a dismiss button. It's an attestation, and clicking it moves the liability from the platform to you. If you don't actually know whether you own what came out, that click is the most expensive thing in this entire workflow.

The fix that cleared it: a woman in a long coat, and a fire escape instead of a leap. Same era, same grade, same three-shot structure. Different silhouette.

Same problem with the Pixar prompt. "Big expressive eyes" plus "exaggerated proportions" plus a child character isn't describing a style, it's describing a specific studio's character design language, which is trade dress rather than copyright and a different kind of problem. So the rewrite specifies rendering instead: global illumination, subsurface scattering, cloth simulation, shallow depth of field. Same look. No house invoked.

Describe the look Soft watercolor, pastel palette, hand-drawn, dreamy edges. Gets you there without typing a studio name.
Not the house Style itself isn't copyrightable in US law. Characters are. Trade dress is its own thing. Stay on the style side.
Watch your ingredients The recognizable creature, the iconic silhouette, the signature prop. These trip filters regardless of your wording.
Disclose your output YouTube auto-detects and labels AI video now. TikTok reads the C2PA tags Seedance embeds. Undisclosed AI risks permanent demonetization.

Worth knowing on that last point: disclosed AI content reportedly earns roughly the same RPM as non-AI content in the same niche. AI is not the disqualifier. Undisclosed AI and low-effort AI are. YouTube's renamed "inauthentic content" policy now targets channels built on mass-produced, templated output, with a three-strike path to permanent removal from the partner program.

And here's the uncomfortable truth about this trend. Its biggest hits are Dragon Ball panels. Somebody's intellectual property. The stuff going most viral is exactly the stuff you can't build a business on. Plan accordingly.

This is practical guidance from published reporting and terms of service, not legal advice. If real money depends on the answer, ask a lawyer rather than an article.

Where to Start

Don't start with a film. Start with one character and one shot.

Generate a face on a white background, no expression, nothing else in frame. Generate an outfit separately. Fuse them into one master sheet. Then write a three-shot prompt against that sheet, one camera move per shot, with a constraints block on the end. Run it at 480p. Look at what you got. Change one phrase. Run it again.

That loop is the whole skill. Everything in this article is an elaboration of it.

Here's the part worth sitting with, though. The reason to learn this isn't that cheap video is valuable. It's the opposite. In China's micro-drama industry, where all of this landed first and hardest, production costs fell by about 90% in ninety days. And of the roughly 128,000 AI dramas in circulation by February, only 0.117% crossed 100 million views. ByteDance's own model team concedes that output quality "does not differ much" between productions, and that standardized prompts converge on the same look.

By April, ByteDance's own Douyin announced a $27.5 million fund to support live-action short drama again. Read that how you like, but the industry read it as an admission the swing had overshot.

Cheap video is not an advantage. It's a commodity input that just got commoditized. Once everyone can make it, the thing that matters is what you make and who you make it for. The workflow in this article won't give you an idea. It'll just stop the idea you already have from falling apart at the render.

So go build the asset sheet. Then go have something to say.