MiniMax H3 can generate synchronized speech, music and environmental sound with video. That is a production feature, not an accessibility layer. A CSR release still needs reviewed captions, a transcript, visual description where needed and operable playback.
Sound can make a clip feel finished. Yet a Deaf viewer may miss the spoken message, a blind viewer may miss meaningful action, and a keyboard user may be unable to control the player. Review the whole release package, not just the generated file.
What native audio does, and what it does not do
MiniMax describes H3 as a multimodal model for four-to-15-second video at 24 frames per second with 32 kHz stereo sound. It reports stable dialogue across 11 languages and says results in additional languages can vary.
These details help a team plan a short scene. They are not an assurance that a pronunciation is locally correct, a social programme is represented accurately, or an accessibility requirement has been met.
| Item in the release package | What it contributes | What still needs review |
| H3 video with native audio | Timed speech, music, ambience and effects | Facts, pronunciation, consent, sound balance and cultural context |
| Captions | A text alternative for speech and meaningful sound | Timing, speaker identification, wording and non-speech cues |
| Transcript | A readable account outside the moving image | Completeness, structure and essential visual information |
| Audio description | Spoken access to important visual action | Selection, timing, clarity and interference with dialogue |
| Media player | Delivery of the video and its alternatives | Keyboard operation, labelled controls and access to captions |
| Human review | Context that automated checks cannot supply | Clear ownership of corrections and final approval |
One row cannot replace another. A transcript does not make an unusable player operable. Captions do not describe a silent chart change, and audio description cannot correct a mispronounced place name.
Start with the communication duty
A CSR video needs one plain purpose before anyone writes a prompt: a safety instruction, a programme result or an application guide, for example. That purpose identifies the information that cannot be lost.
Write an accessibility brief beside the creative brief. It should identify:
- the intended audience and languages;
- the essential facts, names, numbers and call to action;
- meaning carried only by speech, sound, movement, colour or on-screen text;
- the caption, transcript and description deliverables;
- the channel where the clip will appear and the person responsible for each review;
- the source documents against which claims will be checked.
The prompt directs a generation; the brief defines what reviewers must check. A short H3 scene cannot carry every outcome or disclaimer. Put the indispensable message in speech and readable text, provide detail in the surrounding page or transcript, and leave a pause for visual description when needed.
Review a rehearsal before building variants
An early cut exposes errors before they spread across languages and aspect ratios. Use approved copy, since names and numbers are where mistakes matter.
After the accessibility brief is signed off, make a six-second MiniMax H3 rehearsal in ClipDance with every spoken line and meaningful sound written in the production notes. Freeze one candidate export so the language, caption and description reviewers are all examining the same cut.
Then review it in several passes:
- Listen without the picture. Can a listener identify the purpose and required action? Note information found only on screen.
- Watch without sound. Captions should carry dialogue and meaningful cues such as an alarm or speaker change.
- Read the transcript alone. It should work as a document, with headings and necessary visual context.
- Check claims at their source. Verify programme names, dates, quantities, eligibility and contact instructions.
- Use a local-language reviewer. Pronunciation, register and regional usage affect comprehension and trust.
This is a content review, not a model benchmark; its findings apply to that script and export.
Build the alternatives from the approved cut
Accessibility files should follow the locked video, because even a small timing change can break carefully reviewed captions and description.
Captions are more than dialogue
Create captions from the final audio and compare every line with what is heard. Identify unclear speakers and include meaningful non-speech audio, such as a warning tone or machine stopping. Check timing and readability, and keep captions away from on-screen instructions. Translate from approved copy, not an unreviewed automatic transcript.
The transcript should work as a document
A useful transcript is not a block of subtitle text. Add speakers, headings and essential visual information. Make phone numbers and application steps copyable. Keep the transcript beside the video and update it with the cut.
Describe visuals that carry meaning
Audio description covers important visual information absent from the soundtrack: a safety demonstration, service-area map or changing result on a graph. Place it in pauses without competing with essential dialogue. If the main audio already conveys every meaningful visual, record and review the decision not to add a separate track.
Check the route through which people will receive it
The same file can work in one context and fail in another. Test the actual place of publication.
- Can every player control be reached and operated with a keyboard?
- Are play, pause, volume, caption and full-screen controls labelled clearly?
- Can viewers find the intended caption track and nearby transcript?
- Does keyboard focus remain visible?
- If the clip starts automatically, can the viewer stop it promptly?
- Do captions and description work on a phone and slower connection?
Where practicable, test with the audience’s devices and assistive technology. A production-team desktop check cannot substitute for that experience.
Scale only after the package passes review
After the master cut, captions, transcript and description plan pass review, send only the approved generation request through reAPI. Add the returned task ID, exact model ID, prompt version, reviewer and filename to the campaign asset register.
The record supports corrections and withdrawals. It documents production lineage; it does not certify accessibility.
Reopen the relevant checks for every language or edit. Translated text may outgrow a caption window, regenerated action may remove room for description, and a new voice may require different speaker labels.
Assign human review, not just file delivery
Give named reviewers authority to request changes, not merely deliver files.
| Review | A suitable reviewer | The decision to record |
| Factual and programme accuracy | CSR programme owner or subject specialist | Claims match approved records |
| Language and pronunciation | Fluent local-language reviewer | Meaning, names and register are acceptable |
| Caption experience | Caption editor and, where possible, Deaf or hard-of-hearing reviewer | Speech and meaningful sound are available in context |
| Visual access | Audio-description specialist and, where possible, blind or low-vision reviewer | Essential visual information is conveyed clearly |
| Player operation | Accessibility tester using keyboard and relevant assistive technology | Alternatives can be found and controlled |
| Consent and dignity | Safeguarding or community representative appropriate to the project | Representation does not expose or misstate participants |
Budget and schedule participation, then act on it. Keep the issue log with the asset register so corrections have owners.
What this checklist can and cannot establish
This process maps production decisions to W3C media guidance and can expose missing deliverables. It cannot promise legal or WCAG compliance. Requirements depend on jurisdiction, channel, content and implementation. Assess the finished experience against the organisation’s actual duties, especially for public services, employment information and essential instructions.
Frequently asked questions
Does MiniMax H3 native audio remove the need for captions?
No. Native audio means synchronized sound. Captions put dialogue and meaningful non-speech cues into text, and must be reviewed against the final cut.
Are translated subtitles the same as captions?
Not necessarily. Subtitles commonly translate dialogue. Captions also communicate sounds needed for understanding and identify speakers when necessary. A translated deliverable may need both kinds of information.
Does every CSR video require a separate audio-description track?
Not always. If the main audio conveys all essential visuals, added description may be unnecessary. If action, charts or on-screen instructions carry meaning, description or a media alternative may be needed. Review the cut and applicable requirement.
