
Hailuo 3.0 arrived on Runway Dev with an unusually broad reference stack: text, images, video, audio, and keyframe control. That sounds useful, but a feature list cannot tell a production team whether adding a reference protects a shot or introduces a new kind of drift. The practical answer is a controlled ladder in which one story beat stays fixed while references are added one at a time.
Run the Hailuo reference-stack test in APOB
Runway's API changelog states that hailuo3 supports text-, image-, and video-to-video modes, 5–15-second durations, 768P or 2K output, keyframes, and image, video, and audio references. This article treats those as documented capabilities available through Runway's developer platform—not as evidence that APOB currently exposes Hailuo 3.0, and not as a promise about output quality. APOB's current named Hailuo surface is the Hailuo 2.3 AI Video Generator. Evidence checked: September 9, 2026.
The useful question is narrower: when does each extra reference improve control enough to justify its cost and complexity?
Freeze one story beat as the test contract
A reliable Hailuo 3 AI video test begins with a scene that is specific enough to score and simple enough to repeat. Use one subject, one action, one location, one short sound cue, and one endpoint. Save the exact prompt and files before rendering.
Character contract
Describe only observable identity features: age range, hair shape, wardrobe, one prop, and two facial markers. Avoid subjective labels such as “beautiful” or “cinematic.” Create a reference sheet with front, three-quarter, and profile views. Record its pixel dimensions and hash so every rung uses the same image.
Pass identity only when the same markers survive throughout the clip. A plausible stranger is still a failure.
Motion contract
Choose one readable action—for example, a presenter lifts a blue bottle, turns it once, and places it beside a white card. Specify the starting hand, object path, and final position. Do not combine walking, camera orbit, cloth simulation, and product handling in the first test.
Audio contract
Use a short source cue with an obvious onset and ending: a spoken six-word line, a single door close, or a measured music sting. Preserve the original WAV or highest-quality source. The goal is not to judge taste; it is to see whether the requested sound is present, synchronized, and stable.
Acceptance threshold
Set the decision rule before generation. A practical contract might require all identity markers, the complete object action, no extra speech, the cue within a two-frame tolerance, and both endpoint states. These are editorial thresholds for this project, not official Hailuo benchmarks.
Axis | Pass condition | Evidence to save |
|---|---|---|
Identity | All named markers persist | Start, middle, end frames |
Motion | Action order and object path match | Contact sheet or timecodes |
Sound | Intended cue is present and aligned | Waveform screenshot |
Timing | Start, action, and end fit the beat | Frame/timecode log |
Climb the reference ladder one input at a time
Create four rungs: text only, text plus image, then video, then audio. Keep prompt wording, duration, resolution, aspect ratio, and other settings unchanged. Run at least three outputs per rung if budget permits; one lucky result is not a pattern.
Text baseline
The text-only render is the control. It shows what the model invents when it receives no visual or sound evidence. Score it even if it looks attractive. If it violates the contract, record the failure rather than rewriting the prompt; prompt edits would create a different experiment.
Image step
Add the character sheet as the only new input. This Hailuo reference video rung asks whether a still reference improves identity without degrading the action. Note any new wardrobe mutation, background borrowing, or frozen pose. A better face with a broken handoff is a trade, not an unqualified win.
Video step
Add one short motion reference whose action matches the contract. The official Runway entry says reference video is supported and billed by input duration, so keep the clip tight and log its length. Compare pose order, object trajectory, camera motion, and any unwanted style transfer.
Audio step
Add only the prepared cue. Listen with headphones and inspect the waveform. Check whether speech content, onset, duration, and ambient bed match the request. If the audio reference changes picture timing, log that cross-modal effect separately instead of folding it into a vague quality score.
The APOB AI Video Generator can serve as a matched-workflow comparison: reuse the same brief and acceptance rubric, but label the model and date on every output.
Before comparing rungs, check that the uploaded reference files are genuinely identical. Some tools recompress an image or video at upload, so keep the local source, note the interface-reported input, and avoid claiming byte-for-byte preservation unless the platform exposes the resulting file. If an input is rejected or silently omitted, record the rung as incomplete rather than grading its output as though the reference had been used.
Pin the endpoints without scripting every frame
Keyframes are most informative when they constrain the beginning and end while leaving the middle free to reveal the model's behavior. MiniMax's earlier Hailuo 02 start/end-frame announcement demonstrates that control as a product concept; the current Hailuo 3.0 test should still verify the observed result rather than assume lineage guarantees performance.
Start state
Use a clean opening frame that already satisfies the identity, prop, and composition rules. Avoid motion blur and cropped hands. Record whether the generated first visible frame truly matches it or merely approximates the scene.
End state
Build the final frame from the acceptance contract: product on the marked surface, subject looking at camera, card still visible. The endpoint should be achievable within the chosen duration; otherwise the test confuses impossible timing with model drift.
Mid-shot freedom
Do not prescribe every intermediate pose. Let the model solve the transition, then inspect contact points, object continuity, gaze, and acceleration. That middle is where useful motion competence—or hidden interpolation trouble—appears.
Failure capture
Keep the failed clip, settings, and three frames that explain the rejection. Label the failure narrowly: wrong hand, duplicated object, identity swap, missing cue, or endpoint miss. “Bad generation” cannot guide the next test.
Score identity, motion, sound, and timing drift
Use a zero-to-three scale: 3 matches the contract, 2 has a small repairable deviation, 1 has a material deviation, and 0 fails the axis. Score the four axes independently and retain the lowest score; averaging can hide a fatal identity or rights problem.
Ask two reviewers to score independently before discussing the output. Store both raw scores and the agreed result. Disagreement is useful evidence: it often exposes a vague acceptance rule, a frame reviewers inspected at different moments, or a stylistic preference accidentally treated as a defect. Tighten the rubric, not the story of the result.
Identity drift
Compare the same facial markers, clothing boundaries, hand count, and prop identity at three points. Treat flattering changes as drift if they break continuity. Note whether the error begins after adding image, video, or audio.
Motion drift
Track action order, limb path, object contact, and camera behavior. A smooth but reversed action fails the contract. A stiff but correct movement may remain usable if the project permits editing.
Sound drift
Separate missing content, changed words, timing offset, noise, and tonal change. Save a waveform plus a short listening note. Do not claim phase-accurate synchronization unless it has been measured.
Timing drift
Mark the first action frame, peak action, and settled endpoint. Compare them with the planned beat. This makes the AI video drift test reproducible even when two reviewers disagree about style.
Publish the reproducibility card, not a winner claim
The final deliverable is a small evidence card, not “Hailuo 3.0 wins.” It should let another creator repeat the run and understand the limits of the observation.
Input manifest
List prompt, source filenames, hashes, dimensions, durations, and ownership status. Never publish private reference media merely to make the test look complete.
Settings record
Record provider surface, visible model ID, mode, duration, resolution, ratio, seed or job ID when exposed, and generation date. Use the APOB AI Video Prompts page to keep the comparison prompt readable and versioned.
Retry count
Count every submitted generation, not just downloaded favorites. Report accepted outputs per attempt and why each rejection occurred. This is production evidence; it is not a universal success-rate estimate.
Observed boundary
End with the narrowest justified conclusion: which rung improved which axis, what regressed, and under what settings. Recheck on October 9, 2026, or sooner if the provider changes input limits, modes, resolution, or billing.
Add a compact decision row: keep text only, add image, add motion reference, add audio, or hold for repair. The decision should cite the accepted rung, its lowest axis score, total attempts, and known boundary. A future rerun can then compare the same operational choice instead of trying to reconstruct why a team adopted a more complicated stack.
Archive the card beside the tested files and give it a version number. If any source, setting, provider surface, or acceptance rule changes, create a new version rather than replacing the original record.
The best Hailuo keyframes are not the ones that make the boldest demo. They are the ones that survive a fixed contract, a transparent ladder, and a failure log another team can reproduce.
Sources

Be the first to like this.

No credit card needed













