Kuca.ai
Loading
Prompt0 / 50,000
Aspect ratio
Duration5s
5s15s
Resolution
64Credits
Remaining credits
Public

Fast H3 Max generation · four creation modes on Kuca

MiniMax H3 Max: faster generation, precise camera control

MiniMax H3 Max is tuned for creators who want polished video with less waiting. It follows detailed prompts closely, produces visually refined frames, and supports text, image, mixed-reference and camera-control workflows. On Kuca, you can move from brief to preview faster while keeping 480P–1080P output, 5–15 second clips and six ready camera moves in one workspace.

Compare with MiniMax H3
Speed
High-throughput generation
References
9 images · 3 videos · 3 audio
Camera
6 preset trajectories
Output
480P–1080P · 5–15s

H3 or H3 Max? Read the differences before you spend credits

Both models belong to the same H3 generation family, so the choice comes down to what each one was tuned for. H3 Max leans toward prompt adherence, aesthetics and camera precision. MiniMax H3 remains the 2K model with jointly generated stereo sound.

DimensionMiniMax H3MiniMax H3 Max
Tuned for2K detail and native stereo audioPrompt adherence, visual aesthetics and camera precision
Resolution768P, 2K480P, 768P, 1080P
Duration4–15 seconds5–15 seconds
Reference materialUp to 5 images and 3 audio clipsUp to 9 images, 3 videos and 3 audio clips, 12 files in total
Camera controlWritten into the promptSix selectable trajectory presets
Generated audioNative stereo, generated with the pictureNot generated — audio enters as a reference
Prompt lengthUp to 7,000 charactersUp to 50,000 characters
ThroughputBaseline H3 pipelineHigh-throughput pipeline for faster generation
Open the MiniMax H3 page

Both pages open the same generator, so you can move between H3 and H3 Max without rebuilding a brief.

Six camera moves you select instead of describe

Choose one of six ready camera moves for the shot: orbit left, orbit right, dolly in, dolly out, crane up or crane down. Each preset turns camera direction into a clear setting, so your prompt can stay focused on the subject, action and scene.

Orbit left

The camera arcs around the subject, holding it centred while the background slides to the right.

Best forRevealing a product's side profile or the space behind a subject.

Orbit right

The same arc in the opposite direction, with the background sliding to the left.

Best forMatching the direction of existing footage or following on-screen movement.

Dolly in

The camera pushes straight toward the subject and the frame tightens around it.

Best forBuilding emphasis on a face, a label or a piece of detail.

Dolly out

The camera pulls back and the surroundings enter the frame.

Best forClosing a beat by revealing where the subject actually is.

Crane up

The camera rises and the view tilts down onto the subject.

Best forOpening a scene or showing the scale of a location.

Crane down

The camera descends toward the subject's own level.

Best forSettling into an intimate, eye-level finish.

Camera controls take one first frame and one preset, and the prompt is optional in that mode. Aspect ratio follows the frame you upload.

Four endpoints, four different contracts

Each H3 Max endpoint asks for different inputs and leaves different decisions open. Choosing the endpoint first is what keeps a reference set small.

  • 01

    Text to video

    You supply
    A prompt and nothing else.
    The model decides
    Framing, subject, lighting and camera.
    Use it when
    You are still exploring a direction and nothing has to be preserved.
  • 02

    Image to video

    You supply
    A first frame, plus an optional ending frame.
    The model decides
    Motion and timing, while framing follows your image.
    Use it when
    Composition, identity or product geometry has to survive the render.
  • 03

    Mixed reference

    You supply
    Images, video clips and audio clips, each with a named role.
    The model decides
    How those references combine into a single shot.
    Use it when
    Motion, style or sound has to come from existing material rather than a description.
  • 04

    Camera controls

    You supply
    A first frame and one of the six presets.
    The model decides
    Subject behaviour inside the move, because the trajectory comes from the preset.
    Use it when
    The shot exists mainly to show a specific camera movement.

Every reference file should own exactly one attribute

H3 Max accepts a lot of material, but a small set with clear ownership beats a large set that argues with itself. Give each file one job and name that job in the prompt.

AssetWhat it should ownLimitCommon mistake
ImageIdentity, product shape, wardrobe, environment or palette.Up to 9 images.Uploading several near-identical frames that compete for the same attribute.
VideoA camera move, a rhythm, a gesture or the way a material behaves.Up to 3 clips, 2–15 seconds each, 15 seconds combined.Feeding in footage with cuts, text overlays or a different subject.
AudioA voice, a music bed, an ambience or a sound effect.Up to 3 clips, 2–15 seconds each, 15 seconds combined.Expecting the clip to be regenerated: audio is guidance, not output.
PromptSubject, place, action, camera, chronological order and boundaries.Up to 50,000 characters.One vague sentence, or three separate plots squeezed into five seconds.

Images, videos and audio share a combined ceiling of 12 files, and reference mode needs at least one of them.

The exact envelope every request is validated against

These values come from the generator's own configuration. It is quicker to check them while you assemble a shot than after a rejected submission.

Duration
5–15 seconds, whole seconds only
Resolution
480P, 768P or 1080P, defaulting to 768P
Aspect ratio, text to video
21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Aspect ratio, reference mode
The six ratios above, plus adaptive
Aspect ratio, image and camera modes
Follows the first frame you upload
Image references
Up to 9
Video references
Up to 3 clips, 2–15 seconds each, 15 seconds combined
Audio references
Up to 3 clips, 2–15 seconds each, 15 seconds combined
Reference assets in total
12 maximum, and at least 1 in reference mode
Prompt length
Up to 50,000 characters

Camera controls accept a first frame and a preset, and the prompt is optional for that endpoint.

Five failures that waste the most takes

Most disappointing H3 Max renders trace back to the brief or the reference set rather than to the model. These are the patterns worth checking first.

  • The render ignores what the prompt asked for

    Why it happensThe brief is either too vague or contains several unrelated ideas.

    What to changeDescribe one action with concrete nouns, then state the camera behaviour in a separate sentence.

  • The shot feels rushed or cuts off mid-motion

    Why it happensToo much plot is packed into a 5–15 second window.

    What to changeKeep one main action, drop the secondary beat, or extend the duration before adding detail.

  • Text, logos or signage come out warped

    Why it happensA source image already contains lettering, watermarks or busy background clutter.

    What to changeStart from a clean frame, crop the lettering out, or keep on-screen text out of the shot entirely.

  • The subject drifts away from the reference

    Why it happensTwo files are quietly competing for the same attribute, such as identity or colour.

    What to changeGive each attribute a single owner and say so in the prompt, then remove the redundant file.

  • The camera ignores the movement you described

    Why it happensA precise move was written as prose in an endpoint that expects a preset.

    What to changeSwitch to camera controls, choose one of the six trajectories, and let the prompt cover subject behaviour only.

Four shots H3 Max handles well

Each playbook lists what to prepare, how to shape the prompt and what usually goes wrong. Load any prompt straight into the generator and adapt it to your own footage.

  • Vertical social cut

    Prepare
    One clean 9:16 or 3:4 subject image, plus a short reference clip when a specific move matters.
    Prompt shape
    Subject, then place, then a single action, then the camera, then the beat you want in the final second.
    Watch out
    Cramming a full story into five seconds. Keep one action and let it finish.
  • Product turn

    Prepare
    A well-lit product image on a plain surface, ideally without printed text or reflective clutter.
    Prompt shape
    Product, then surface, then the camera preset, then notes on lighting and material.
    Watch out
    Lettering on packaging is the first thing to warp. Crop it out or accept a softer logo.
  • Open-world action beat

    Prepare
    Character and environment images, plus a short clip that demonstrates the movement you want.
    Prompt shape
    Character, then environment, then one action, then the camera, then continuity boundaries.
    Watch out
    Reusing recognisable game or film assets invites IP trouble. Keep every design original.
  • Trailer-style teaser

    Prepare
    A hero image and a music reference, rendered as several short takes instead of one long shot.
    Prompt shape
    A shot list of 5–15 second beats, each with its own camera preset and a single action.
    Watch out
    One request cannot carry a whole trailer. Build it take by take and cut it together in an editor.

Keep source material original or properly licensed. Do not upload game assets, film frames, logos or other protected characters and brands, and avoid prompts that imitate a specific franchise.

Questions people ask before their first H3 Max render

MiniMax H3 Max is a faster, more controllable version of the H3 video model. It improves prompt adherence and visual polish while adding mixed image, video and audio references plus six ready camera moves.


Two qualities improved: prompt adherence, meaning how faithfully the render follows a written brief, and visual aesthetics, meaning how the finished frame is lit, composed and graded. H3 Max is not a larger model trained purely for speed.


H3 Max is built for higher-throughput inference, so it can process a shot faster than the baseline H3 workflow. Actual completion time still varies with resolution, duration, reference files and queue load.


H3 Max responds to prompts more closely, accepts video references alongside images and audio, and offers six preset camera trajectories. MiniMax H3 stays the better choice for 2K output and for native stereo audio generated together with the picture.


480P, 768P and 1080P, with 768P as the default, and any whole-second duration from 5 to 15 seconds. H3 Max does not offer the 2K option available on MiniMax H3.


Up to 9 images, 3 video clips and 3 audio clips, capped at 12 files in total. Video and audio clips each run 2–15 seconds, with a 15-second combined budget per category, and reference mode requires at least one file.


Kuca provides six selectable moves: orbit left, orbit right, dolly in, dolly out, crane up and crane down. Camera control takes a first frame plus one preset, and the prompt becomes optional.


No. In Kuca, audio is an input: up to 3 reference clips of 2–15 seconds that can pin a voice, a music bed, an ambience or a sound effect. If you need a soundtrack generated with the picture, use MiniMax H3.


Work through six parts in order: subject, place, action, camera, chronological order and boundaries such as what must not appear. Then keep one main action per 5–15 second shot and change a single variable between takes.


Vague prompts, several plots packed into a short duration, source images that already contain text or distracting clutter, and reference files that compete for the same attribute. The diagnostics section on this page covers each one and what to change.


Start an H3 Max shot

Choose a mode, add your brief or references, and start a faster H3 Max generation directly on Kuca.