Seedance 2.5 vs. MiniMax H3: What's The Difference?

Explore all models

Both models write audio jointly with the picture, so dialogue and room tone arrive already timed to the cut. Seedance 2.5 runs 4 to 30 seconds at 480p, 720p or 1080p and bills on tokens, at $0.0214 per 1,000 up to 720p and roughly $0.0234 at 1080p. MiniMax H3 runs 5 to 15 seconds at 480P, 768P, 2K or 4K and bills per second, from $0.05 up to $0.16. Seedance 2.5 doubles the runtime out of one generation, while MiniMax H3 delivers above 1080p at under half the cost per second.

last updated
8/19/2026
edited by
John Ozuysal
read time
15 minutes
Seedance 2.5 vs. MiniMax H3: What's The Difference?

In this guide, I'll put both models through the same eight shots on fal, then go through the specs and the pricing that decide which one you point at a given job.

TL;DR

Both models write audio jointly with the picture, so dialogue and room tone arrive already timed to the cut, with no second pass.

Seedance 2.5 runs 4 to 30 seconds at 480p, 720p or 1080p and bills on tokens, at $0.0214 per 1,000 up to 720p and roughly $0.0234 at 1080p.

MiniMax H3 runs 5 to 15 seconds at 480P, 768P, 2K or 4K and bills per second of output, from $0.05 up to $0.16.

Duration and resolution pull in opposite directions here: Seedance 2.5 doubles the runtime you can get out of one generation (up to 30 seconds!), while MiniMax H3 delivers above 1080p at under half the cost per second ($0.16 per second at 4K).

How research was conducted: all eight clips ran at 10 seconds on fal, Seedance 2.5 at 1080p and MiniMax H3 at 4K.

Prompt text is identical on both models except for the reference tokens, which each model spells differently.

MiniMax H3's prompt expansion was left enabled, which is its default.

How does Seedance 2.5 compare to MiniMax H3?

Here's how both models compare head-to-head:

Seedance 2.5MiniMax H3
DeveloperByteDanceMiniMax
WeightsClosedOpen, published August 2026
Best forSingle takes past 15 seconds, large reference packs2K and 4K delivery, lower cost per second
Resolutions480p, 720p, 1080p480P and 768P native, 2K and 4K upscaled from a 768P base
Duration4 to 30 seconds, or auto5 to 15 seconds
Aspect ratiosauto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:1621:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive on reference
Billing unitPer 1,000 tokens: $0.0214 to 720p, ~$0.0234 at 1080pPer second of output
Lowest tier~$0.2205 per second at 480p$0.05 per second at 480P
Mid tier$0.4730 per second at 720p as published, $0.4622 by the formula$0.13 per second at 2K
Highest tier~$1.14 per second at 1080p by the formula$0.16 per second at 4K
Frame area affects price✅ Yes, it feeds the token count❌ No, the per-second rate is flat
Native audio✅ Written with the picture✅ Native stereo, written with the picture
Audio togglegenerate_audio, on by default❌ Not exposed, audio comes back either way
Lip-sync✅ Dialogue in double quotes✅ Dialogue generated with the picture
Reference audio, documented useRhythm, timing and voiceCited by order in the prompt; fal's model page also describes voice transfer
End frame control✅ Yes✅ Yes
Prompt expansion❌ Not exposed✅ On by default
Safety checker toggle❌ Not exposed✅ API only, locked on in the playground
EndpointsText, image and reference to videoText, image and reference to video
Reference files, combined cap5012
Reference imagesUp to 30Up to 9
Reference videosUp to 10, each 1.8 to 30.2 secondsUp to 3, each 2 to 15 seconds
Reference audioUp to 10, each 1.8 to 30.2 secondsUp to 3, each 2 to 15 seconds
Reference addressing@Image1, @Video1, @Audio1Image 1, Video 1, Audio 1
Reference billingInput video seconds billed, then a 0.6 multiplierFirst 5 images free, $0.08 per image after

Where can you access Seedance 2.5 and MiniMax H3?

The best place to access Seedance 2.5 and MiniMax H3 is fal, as there's only one API key to manage and no capacity to reserve on your side.

Both models are in the fal playground and on the API, and the @fal-ai/client SDK treats them the same way it treats every other video model on the platform, so nothing about queueing, webhooks or error handling changes when you swap one for the other.

Here's what it looks like calling Seedance 2.5:

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("bytedance/seedance-2.5/text-to-video", {
  input: {
    prompt:
      "An octopus finds a football in the ocean and excitedly calls its octopus friends to come and play. Cut scene to an octopus football game under the sea.",
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});

console.log(result.data);
console.log(result.requestId);

And how it looks when calling MiniMax H3:

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("minimax/h3/text-to-video", {
  input: {
    prompt:
      "A white kitten chases a butterfly across a sunlit garden. Gentle camera tracking, natural movement, soft afternoon light filtering through the leaves.",
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});

console.log(result.data);
console.log(result.requestId);

Seedance 2.5 vs. MiniMax H3: text-to-video tests

To show you how both models perform, I'll go over four shots on their text-to-video endpoints, each pointed at a different failure mode.

Test 1: A rescue swimmer going into heavy sea (21:9)

Prompt: Ultrawide 21:9, filmed from a second aircraft holding station off the port side. Last light, heavy overcast. [0 to 4 seconds] A search and rescue helicopter holds a low hover over grey swell, nose down a few degrees, rotor wash flattening a bright disc into the water beneath it. A rescue swimmer braces in the open door, legs outside the skid, both hands on the frame. Camera level with the cabin, matching the hover. Sound: blade slap over everything, turbine whine behind it, wind across the mic. [4 to 7 seconds] He pushes off and drops feet first, entering clean between two crests. The disc under the aircraft holds its shape while he clears it. Camera follows him down, arriving a beat behind. [7 to 10 seconds] He surfaces and swims out of the wash toward an orange marker on the next crest. The helicopter climbs away to the left. Sound: sea noise rising as the rotor recedes, breathing over it. No music, no dialogue. Grade: cold and desaturated, spray on the front element, long lens compression, no lift in the blacks.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

💡 Check out our MiniMax prompt guide as well as Seedance 2.5 prompting guide.

Test 2: Three animated characters, one conversation (9:16)

Prompt: Format: stop-motion clay animation, feature quality, vertical 9:16. Stepped motion at twelve frames per second, no motion blur on the puppets. Set: the lit cabin of a night bus, rain on the windows, warm interior against cold blue glass. Cast: BADGER, a conductor in a brass-buttoned coat, aisle seat, thumbprints visible in the clay of his face. HERON, a courier in a nylon jacket, window seat, parcel on her lap, beak hinged at the jaw. TORTOISE, asleep across the back seat. Mouth stays shut for the whole clip. Action: 0 to 4s. The badger punches a ticket, looks up and says: "That parcel has been on this route four nights running." 4 to 7.5s. The heron keeps her eyes on the window. She answers: "It gets off when someone signs for it." 7.5 to 10s. Silence. The bus takes a corner, the parcel slides an inch, and she puts a wing on it. Camera closes on the pair. Constraint: one mouth in motion at a time. The other two puppets are locked still while a line is running. Sound: the badger low and gravelly, the heron clipped and quiet. Rain on glass, diesel underneath, one ticket punch, the parcel sliding. Nothing scored.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

Test 3: A dive watch and a title card (16:9)

Product video usually breaks at the hands, so I check the bezel turn before anything else.

A prompt-writing note that cost me a take on Seedance 2.5: keep on-screen type outside double quotes.

The model reads quoted text as a spoken line, so a title card written the obvious way comes back with a voice reading the brand name over it.

Prompt: Four shots, 16:9, high-end watch commercial. Black on black, controlled speculars, everything outside the key falling away. (0 to 2.5s) A steel dive watch, black ceramic bezel, matte dial, face up on wet basalt in a dark studio. Seawater sheeting off the case. One hard key from the upper left splits the polished bevels from the brushed flanks. Camera starts tight on the crown and pulls back into a three-quarter. (2.5 to 5.5s) A hand comes in from the right and turns the bezel counterclockwise a quarter rotation. Six clicks, evenly spaced. Beads break and run as it moves. The hand leaves the way it came. (5.5 to 8s) The key narrows to a slot across the dial. The seconds hand sweeps the sub-dial once. A droplet on the crystal catches the light and falls. (8 to 10s) Four frames of black, then two lines of white serif type build in centered, letter by letter, wide-tracked: MERIDIAN 300 above, AUTOMATIC DIVER below. Both lines complete before the clip ends. Sound: deep room tone, six evenly spaced bezel clicks, the droplet landing on stone, a low tone rising under the type. Nothing spoken.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

💡 Check out some of the Seedance 2.5 workflows that you can run with the model.

Test 4: A pit stop in four cuts (16:9)

Written timings are usually easy to ignore, and a model that treats them as suggestions will still hand you something that looks fine until you check it against the prompt.

Let's see how the 2 best-in-class video generators will deal with this prompt:

Prompt: Ten seconds, four shots, floodlit pit lane at night. Hard cuts throughout, no dissolves, no fades. Shot 1, 2.5 seconds. Wide from across the lane. A prototype race car in its box under a bank of white floods, crew over it, tire trolleys and air lines on the ground, heat coming off the bodywork. Shot 2, 2.5 seconds. Low at the left front corner. Gun on, nut off, old wheel back, new wheel on, the whole change inside this one shot. Backlit hard enough that the crew reads as silhouette. Shot 3, 2.5 seconds. Macro on the right rear as it takes the car's weight. Grit picking up off the concrete, sidewall flexing. Shot 4, 2.5 seconds. The wide from shot 1 again, same camera position. Box empty, one air line still moving on the ground, floods burning into bare concrete. Camera: handheld, light drift, cold floods against warm sodium spill from the paddock. Heavy haze. Practicals only. Sound: air line hiss, the gun in two bursts, an engine idling then pulling away, crew over the top of it, a distant public address. Nothing scored.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

falMODEL APIs

The fastest, cheapest and most reliable way to run genAI models. 1 API, 100s of models

falSERVERLESS

Scale custom models and apps to thousands of GPUs instantly

falCOMPUTE

A fully controlled GPU cloud for enterprise AI training + research

Seedance 2.5 vs. MiniMax H3: image-to-video tests

Your still image becomes frame one on both models, and both take a second image to land the clip on a specific composition. Output aspect ratio follows the input image either way.

Test 5: A coat that has to move like wool (16:9)

Heavy wool is the whole test. It should swing with weight, flare late, and settle back down.

Editing the end frame off the start frame fixes the face and the coat at both ends, which leaves motion and lighting as the only things either model has to invent.

Start frame, generated with Seedream 5.0 Pro: A photoreal 16:9 product still of a heavy charcoal wool coat on a headless linen tailor's dress form, mounted on a low black turntable on a grey studio sweep. Deep knife pleats falling from a wide buckled belt, brushed steel post and base below the hem. One hard key from the left, everything else in shadow, pleats hanging dead straight. No text, no logos, no props, no people.

Generated using Seedream 5.0 Pro on fal, an AI model from ByteDance.

End frame, produced by running Seedream 5.0 Lite Edit over the start frame: Preserve the coat, the belt, the dress form, the turntable, the grey sweep and the camera position without alteration. Rotate the form forty-five degrees to its left so the coat is turned away from camera, at rest, pleats hanging straight again after the turn, belt still buckled. Add two more heads so the key softens and a rim separates the shoulder line from the background. No text, no logos, no props, no people.

Edited using Seedream 5.0 Lite Edit on fal, an AI model from ByteDance.

Video prompt: The turntable holds for a beat, then carries the dress form left through roughly three-quarters of a turn and stops with the coat angled forty-five degrees away from camera. The pleats lift and flare as the rotation builds, hold their weight through the top of the arc, and drop back against the form over the final second without floating. The belt stays buckled and the hem never lifts above the knee line. Two additional heads come up during the turn, softening the key and striking a rim across the shoulder line. Locked wide, no camera movement at all. Nobody enters frame at any point. Sound: large studio room tone, wool moving against itself, the turntable motor under it, a lighting rig ticking as the second head warms. Nothing spoken, nothing scored.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

Test 6: A puppet that has to stay a puppet (9:16)

Stop-motion is harder to hold than a painted still, because the giveaway is in the cadence and not the surface.

Twelve frames per second with no motion blur is a strange ask for a model trained mostly on smooth footage.

The steam pass is the second check: the fox goes behind it in one pose and has to come out in the same one.

Source image, generated with GPT Image 2: A 9:16 vertical still of a handmade stop-motion puppet set photographed on a real tabletop. A felt-and-wire fox in a knitted scarf stands on a matchbox-scale railway platform built from balsa and card, under a working brass lamp the size of a thimble. Visible felt fibers, wire armature showing at one wrist, dried glue at the balsa joints, dust hanging in the light. Shallow depth of field, warm practical on the platform, cold blue falling away past its edge. No text, no people.

Generated using GPT Image 2 on fal, an AI model from OpenAI.

Video prompt: The fox turns his head down the track, raises one arm to check a pocket watch, lowers it. Steam drifts across from the right and covers him for roughly a second. He emerges in the pose he went in with, standing where he was standing, scarf hanging as it hung. The lamp flickers once. Slow push toward the platform. Throughout: twelve frames per second, stepped, no motion blur on the puppet, felt fibers readable, wire visible at the wrist, glue still on the balsa. The image must not resolve toward smooth CG or live action at any point. Sound: low hall reverb, the watch lid, steam venting, a station bell some distance off. Nothing scored.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

Seedance 2.5 vs. MiniMax H3: reference-to-video tests

In this section, we'll see how Seedance 2.5 and MiniMax perform head-to-head with identical reference-to-video tests, where we supply them with the same references:

Test 7: Three images composited into one shot (16:9)

Three references, three jobs, and one detail small enough to lose: a hairline crack across the faceplate glass in @Image1.

@Image1, generated with GPT Image 2: A 16:9 photoreal studio still of a weathered brass diving helmet, three-quarter view on a wooden stand. Green verdigris worked into the seams, a dented crown, one hairline crack running left to right across the front faceplate. Raking light from camera right, background unlit. Nothing else in frame. No text, no branding.

Generated using GPT Image 2 on fal, an AI model from OpenAI.

@Image2, generated with Seedream 5.0 Pro: A 16:9 photoreal shot inside a flooded stone crypt beneath a cathedral. Groin-vaulted ceiling, still water at knee height mirroring the vaults, one collapsed section dropping a shaft of daylight onto the surface, salt bloom creeping up the pillars. No people, no text.

Generated using Seedream 5.0 Pro on fal, an AI model from ByteDance.

@Image3, generated with Seedream 5.0 Pro: A 16:9 photoreal still of coiled hemp rope and cast-iron ballast weights on a rack against a bare wall. Frayed rope ends, rust bleeding down the wall beneath the weights. Soft even light against a flat grey backing. No text, no branding.

Generated using Seedream 5.0 Pro on fal, an AI model from ByteDance.

Prompt: @Image1 is the helmet. @Image2 is the crypt. @Image3 is the rope and weights. Carry the helmet across intact: brass patina, verdigris in the seams, the dented crown, the hairline crack. The crypt stays as it was photographed, water level unchanged, vaults unchanged, daylight shaft unchanged. Take only the rope and the iron from @Image3 and discard everything around them, including the wall and the grey backing. [0 to 4 seconds] Wide, shooting the length of the crypt. A stone plinth stands just clear of the waterline with the helmet on it, lit by the daylight shaft alone. Rope coiled at the base, two weights half submerged alongside. [4 to 7 seconds] Cloud crosses outside and the shaft dies. A work lamp on a stand behind camera strikes cold and hard, and the helmet's reflection resolves in the water. [7 to 10 seconds] Camera travels down the crypt at water level, finishing close on the faceplate with the crack lit by the lamp. Sound: long stone reverb, water coming down from the vaults, a ballast humming up as the lamp strikes, rope creaking as it settles. Nothing scored. Photoreal throughout. Decay as found, nothing added to the set.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

Test 8: Video, image and audio in one call (16:9)

Three references doing three unrelated things:

@Video1 carries the cutting and the camera.

@Image1 carries the car.

@Audio1 carries a voice belonging to nobody on screen.

@Image1, generated with GPT Image 2: A 16:9 photoreal still of a fictional two-seat electric prototype coupe. Matte dark blue, exposed carbon front splitter, black centre-lock wheels, a single yellow accent stripe along the sill. Three-quarter front view on bare concrete, soft even light, flat grey backing. Unbadged, unbranded, no text, nobody in shot.

Generated using GPT Image 2 on fal, an AI model from OpenAI.

@Audio1, generated with Seed Audio 1.0: Single English speaker, dry studio, dead room, no ambience and no second voice. A man in his fifties, level and unhurried, reading a note and not selling anything. Roughly ten seconds, with a long beat before the final sentence so it can land late: "The stop is eleven seconds if nothing goes wrong. Nothing goes wrong about half the time."

Generated using Seed Audio on fal, an AI model from ByteDance.

Prompt: @Video1 supplies the cutting and the camera. @Image1 supplies the car. @Audio1 supplies the voiceover. Match @Video1 shot for shot. Same four durations, same hard cuts at the same points, same handheld drift, and shot four returns to the shot one camera position exactly as it does in the reference. The car from @Image1 occupies the pit box in place of the car in @Video1. Matte dark blue, exposed carbon splitter, black centre-lock wheels and the yellow sill stripe hold across all four shots without variation. @Audio1 plays across the full ten seconds as voiceover. Nobody in shot mouths any of it and nobody acknowledges the lens. Land the final sentence on the fourth cut. The air line hiss, the wheel gun, the engine and the crowd from @Video1 stay in the mix underneath the voice.

Generated using Seedance 2.5 on fal, an AI model from ByteDance.

Generated using MiniMax H3 on fal, an AI model from MiniMax.

How do Seedance 2.5 and MiniMax H3 compare on pricing?

Seedance 2.5 bills tokens, and MiniMax H3 bills seconds, so nothing lines up until you fix a clip length.

tokens = (output height x output width x duration x 24) / 1024

EndpointSeedance 2.5MiniMax H3
Text to video$0.0214 per 1,000 tokens to 720p, roughly $0.0234 at 1080pPer second of output: $0.05 at 480P, $0.08 at 768P, $0.13 at 2K, $0.16 at 4K
Image to videoSame token rates, input stills not billedSame per-second rates, input stills not billed
Reference to videoSame token rates, with input video seconds added to the formula and the total then multiplied by 0.6. Image and audio references freeSame per-second rates. First 5 reference images free, $0.08 each after. Reference video billing not addressed in fal's note

💡 Rates published on fal as of the 16th of August, 2026.

Which one should you use: Seedance 2.5 or MiniMax H3?

I wouldn't really say that there is one obviously superior AI video generator and the purpose of this article wasn't really to select a "winner".

However, here's how I see the differences between the models:

Choose Seedance 2.5 if

The take has to run past 15 seconds in one piece: 30 native seconds covers a product demo that needs a setup before its payoff, or a testimonial allowed to finish its thought, with no splice to hide.

You are feeding it more than a dozen references: 50 files against 12 is the gap between holding one character through one clip and holding a cast and a standing set across a whole series.

You need a switch for silence: generate_audio defaults to on and can be turned off, which matters when picture is going to someone else's mix.

1080p is the delivery spec, and you want it generated at that size: fal gives Seedance 2.5's 1080p its own token rate and describes no upscaling step.

Choose MiniMax H3 if

Cost per second is the binding constraint: ten seconds at 4K runs $1.60, against $4.62 for ten seconds of 720p and about $11.37 at 1080p on Seedance 2.5.

The spec says 4K: nothing else here reaches it. Read the mechanism though: fal documents 2K and 4K as upscales of a 768P generation.

Local deployment could come up later: MiniMax published the weights in August 2026.

💡 Neither one needs a subscription on fal. You pay for what you successfully generate and nothing else.

Recently Added

Get started with Seedance 2.5 and MiniMax H3 on fal

Both Seedance 2.5 and MiniMax H3 are live on fal today, and they can be accessed through our playground, API, Agent, Sandbox, or MCP, with one API key covering all endpoints.

Seedance 2.5 runs at bytedance/seedance-2.5/text-to-video, bytedance/seedance-2.5/image-to-video and bytedance/seedance-2.5/reference-to-video.

MiniMax H3 runs at minimax/h3/text-to-video, minimax/h3/image-to-video and minimax/h3/reference-to-video.

Get started for free at fal.

Seedance 2.5 vs. MiniMax H3 FAQs

What's the main difference between Seedance 2.5 and MiniMax H3?

Both are omni-modal models that read text, images, video and audio as one context and generate stereo sound jointly with the picture.

Seedance 2.5 goes to 1080p, 30 seconds in a single continuous pass, and up to 50 reference files.

MiniMax H3 trades runtime for resolution and cost, generating natively at 768P and upscaling that base to 2K or 4K, across clips of 5 to 15 seconds and a combined cap of 12 reference files.

MiniMax also published H3's weights, which Seedance 2.5 does not offer.

Which one is cheaper?

MiniMax H3, at every resolution on offer.

Ten seconds costs $1.60 at 4K and $0.80 at 768P, against $4.62 for ten seconds of 720p on Seedance 2.5 and roughly $11.37 at 1080p.

Reference work narrows the gap a little, since Seedance 2.5 applies a 0.6 multiplier there, though input video seconds get billed on top.

Do both models generate audio?

Yes, on all six endpoints, and neither model charges extra for it.

Dialogue, effects and ambience are written into the same pass that makes the picture.

Seedance 2.5 exposes a generate_audio flag that defaults to on, which you can switch off when you want silence. MiniMax H3 has no equivalent toggle and returns audio either way.

Can MiniMax H3 generate 4K?

Yes, on all three endpoints at $0.16 per second, produced by upscaling a 768P base result.

Seedance 2.5 tops out at 1080p on fal, billed at its own token rate and not documented as an upscale.

Can I use both video generators in the same project?

Yes, and it's a common setup.

Both Seedance 2.5 and MiniMax H3 run on fal behind the same @fal-ai/client SDK, so switching between them is an endpoint string and a couple of parameter names.

about the author
John Ozuysal
Founder of House of Growth. 2x entrepreneur, 1x exit, mentor at 500, Plug and Play, and Techstars.

Related articles