
How Much Faster Does the MiniMax H3 Speed-Up LoRA Actually Make Things? A 20-Video Comparison with FL2 and REF4
Compared FL2's arry v4_step600 and REF4's PDD Ref2VA with and without the speed-up LoRA across 20 videos, measuring generation times and visual differences.
I added the speed-up LoRA to MiniMax H3's FL2 and REF4, then tested how much it changed generation time under identical conditions.
Here is the setup I used.
- FL2:
arry v4_step600. The standard version used 20 steps; the speed-up LoRA version used 8 steps. - REF4:
PDD Ref2VA. The standard version used 20 steps; the speed-up LoRA version used 8 steps. - All videos were 8 seconds, 768×1344, at 24 fps.
- Each pair used the same prompt, seed, and reference image.
- The left side is without the speed-up LoRA; the right side is with it.
The tests covered five scenarios: fast motion, intense scene transitions, nearly static footage, fine hair and fabric detail, and camera movement with foreground occlusion. I made 10 videos with FL2 and 10 with REF4, for 20 in total.
FL2 — arry v4_step600
1. Fast motion
A parkour video set in a rainy neon city. The conditions included fast movement, water splashes, and plenty of clothing and hair motion.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 504 seconds (8 min 24 sec); with it, 272 seconds (4 min 32 sec). The non-LoRA result includes the initial time needed to load FL2 into VRAM. When switching to the LoRA version, the container connection dropped once, so the time shown is from the successful retry.
2. Intense scene transitions
Keeping a chrome hummingbird fixed, the video moved in rapid succession through an art museum, jungle, blast furnace, ice cave, paper city, and outer space.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 454 seconds (7 min 34 sec); with it, 236 seconds (3 min 56 sec). In this pair, the LoRA version was about 1.92× faster.
3. Minimal-change video
A locked-off shot of a watchmaker. Only blinking, millimeter-scale tweezer movement, the second hand, and dust were animated.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 447 seconds (7 min 27 sec); with it, 233 seconds (3 min 53 sec). Generation time barely changed even with less motion, for a difference of about 1.92×.
4. Fine hair and fabric detail
A dancer spins rapidly in a translucent chiffon dress embroidered with silver thread. This scenario was meant to see how much fabric and hair would break down at higher speed.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 448 seconds (7 min 28 sec); with it, 234 seconds (3 min 54 sec). This was also about 1.91× faster.
5. Camera movement and foreground occlusion
A mountain bike riding through a forest. Ferns completely blocked the frame while the camera moved around by 180 degrees.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 448 seconds (7 min 28 sec); with it, 234 seconds (3 min 54 sec). Even with extensive camera movement, it was about 1.91× faster.
Across the four FL2 cases after VRAM was loaded, the average was 449.25 seconds without the LoRA and 234.25 seconds with it. In this environment, that works out to about 1.92× faster.
REF4 — PDD Ref2VA
For REF4, I supplied a six-panel storyboard made with ChatGPT image generation as a single reference image.
1. Fast motion
I turned the same parkour sequence as FL2 into video using six pose panels.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 481 seconds (8 min 1 sec); with it, 299 seconds (4 min 59 sec). The non-LoRA result includes the time needed to load REF4 into VRAM. The first LoRA attempt disconnected during decoding, so the reported time is from the successful retry.
2. Intense scene transitions
The same hummingbird and six worlds were developed from a six-panel storyboard.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 424 seconds (7 min 4 sec); with it, 251 seconds (4 min 11 sec). It was about 1.69× faster.
3. Minimal-change video
I turned the watchmaker's subtle changes into six panels. With the speed-up LoRA, the storyboard panel layout remained visibly present in the resulting video.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 418 seconds (6 min 58 sec); with it, 247 seconds (4 min 7 sec). It was about 1.69× faster, but the visual difference was much more pronounced in this case.
4. Fine hair and fabric detail
The same dancer's spin was referenced through six panels. Both with and without the LoRA, this expanded into a relatively clean full-screen video.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 420 seconds (7 min 0 sec); with it, 249 seconds (4 min 9 sec). It was about 1.69× faster.
5. Camera movement and foreground occlusion
A six-panel sequence specified the mountain bike approaching, a foreground wipe, and a move around to the rear. This time, the panel layout remained fairly prominent in both versions.
| Without Speed-Up LoRA · 20 steps | With Speed-Up LoRA · 8 steps |
|---|---|
再生 |
再生 |
Without the speed-up LoRA, it took 419 seconds (6 min 59 sec); with it, 248 seconds (4 min 8 sec). It was about 1.69× faster.
Across the four REF4 cases after VRAM was loaded, the average was 420.25 seconds without the LoRA and 248.75 seconds with it. In this environment, that works out to about 1.69× faster.
What I learned
Looking at generation time alone, the speed-up LoRA has a substantial effect. FL2 was nearly twice as fast, while REF4 cut generation time by roughly 40%. It looks especially useful for iterative testing while refining prompts and compositions.
That said, when REF4 receives a single storyboard contact sheet, it can sometimes treat the grid itself as part of the composition. This was particularly noticeable with the speed-up LoRA, where panel divisions persisted longer in some cases. If you want a normal full-screen video with REF4, providing each panel as a separate image seems safer than using a single contact sheet.
Also, with both FL2 and REF4, the container connection dropped once only when switching from the standard version to the speed-up LoRA version for the first time. Re-running under the same conditions succeeded, and subsequent runs were stable. It would probably be best to generate one short test video after the initial switch.
These figures are measurements from this machine, at this resolution, for these 8-second videos. They should be viewed less as absolute benchmarks and more as a comparison between LoRA and non-LoRA runs in the same environment.























