Grok Imagine Video 1.5 Adds Face and Voice Locking
xAI's Imagine Video 1.5, updated on 31 July 2026, locks a character's face and voice across scenes using up to seven references, plus native 1080p output.
xAI updated Imagine Video 1.5 on 31 July 2026 with a reference system that lets users lock specific elements in place across a generation — most significantly, a character's face and voice together. The company's description is direct: "pass in a character image and a voice reference, and both hold," producing "the same face and the same voice in every scene" [1]. The model accepts up to seven references per generation, and xAI says users can hold a face, product or location constant while varying everything else, or invert that and keep the scene fixed while changing what appears in it [1].
The same update brought native 1080p output to both text-to-video and image-to-video generation, and what xAI describes as "better audio" than the previous version [1]. The company also notes that its text-to-video pipeline "pairs our image generation with image-to-video," meaning a text prompt is rendered as a still first and then animated [1].
What the References feature actually does
Character consistency is the practical wall most generative video work hits. A model can produce an excellent five-second clip, but producing a second clip in which the same person appears — recognisably the same, with the same face, clothing and voice — has generally required either heavy manual intervention or abandoning generative video for the shot. xAI's approach is to let users supply explicit reference inputs that the model treats as fixed rather than suggestive.
Seven references, and what that allows
The seven-reference ceiling is higher than most comparable reference systems, and matters because real production shots rarely involve just one fixed element. A product demo might need a consistent presenter, a consistent product and a consistent location simultaneously, with only the camera angle and action changing between clips. xAI's framing suggests the system is designed for exactly that kind of combination rather than single-subject consistency alone.
Voice as a reference is the less obvious part
Locking a face across scenes is a capability several video models have pursued. Locking a *voice* alongside it, within the same generation request, is less common — most workflows handle visuals in a video model and audio separately in a voice tool such as ElevenLabs, then synchronise the two afterwards. If xAI's implementation works as described, it removes a synchronisation step rather than just improving output quality. That said, xAI has published no sample-based evaluation of how well voice references hold across longer or more varied generations, so this remains a company claim.
Availability is more restricted than the headline suggests
The rollout was staged, and the details matter. According to xAI, image and voice references launched in the United States only, for SuperGrok Heavy and SuperGrok Plus subscribers, on grok.com/imagine and iOS, and were described as "rolling out to all tiers over the next few days" from 31 July 2026 [1]. Text-to-video and native 1080p, by contrast, were made generally available across grok.com/imagine, iOS and Android [1]. On the API, image references and text-to-video are available directly, while voice reference support is listed as "available on request" rather than openly accessible [1].
xAI did not publish pricing for the model or its reference features in the announcement.
Why this matters
Generative video has spent two years producing impressive individual clips that could not be assembled into anything longer, because nothing stayed consistent between them. Reference systems are the industry's answer to that, and the specific combination xAI is offering — face and voice held together, up to seven fixed elements, native 1080p — targets the assembly problem rather than the single-clip quality problem. It arrives in a busy period for the category: HeyGen shipped a batch of agentic video features in July 2026, and Adobe extended Firefly into audio generation on 20 August 2026, both pointing at the same underlying goal of producing finished video rather than raw clips.
Who should care
Marketing teams producing multi-shot product or explainer videos are the most direct beneficiaries, since consistency across shots is precisely what makes a sequence usable rather than a collection of takes. Existing SuperGrok Heavy and Plus subscribers in the US had first access and can test the reference system now. Developers building video generation into a product should note the split availability: image references and text-to-video are API-accessible, but voice references require a request to xAI, which is a meaningful constraint if voice consistency is the reason you are evaluating the model at all.
Practical implications for buyers and users
The right test is a multi-shot sequence, not a single generation. Produce three or four clips of the same character in different scenes and check whether a viewer would accept them as the same person — that is the standard the feature is claiming to meet. Teams outside the United States should confirm current availability before planning around the reference features, since the initial launch was US-only and xAI's "next few days" expansion timeline was stated on 31 July 2026 without a confirmed completion date. For anyone budgeting a project around this, the absence of published pricing means cost per finished minute cannot yet be estimated from public information.
Limitations, availability and unresolved questions
xAI has not published maximum clip duration for Imagine Video 1.5, which is arguably the second most important specification after consistency and is absent from the announcement. There is no published guidance on likeness rights — specifically, whether users may supply a real person's face and voice as references, and what verification or consent requirements apply, which is a significant question for a feature designed to reproduce a specific person across scenes. Commercial usage terms for generated video are not stated in the announcement. It also remains unclear whether the API's "on request" voice reference access is a capacity limitation, a safety gate, or a commercial one.
Verdict
The reference system addresses the correct problem. Character and voice consistency is the specific barrier between generative video as a novelty and generative video as a production tool, and holding both together with up to seven fixed elements is a more ambitious approach than most competing implementations. What is missing is the surrounding detail a business needs: no published duration limits, no pricing, and no stated policy on using real people's faces and voices as references — a notable gap for a feature whose entire purpose is reproducing a specific individual consistently. Worth testing now if you already hold a qualifying subscription; harder to plan a workflow around until xAI publishes the terms.
The current AI Video & Avatars shortlist
Where this sits in the wider market: our current shortlist for AI Video & Avatars, what each tool is best at and the main caution to check before committing.
| Tool | Best for | Current position | Important caution |
|---|---|---|---|
| Runway Production pick | Cinematic generation and controlled editing | Seedance 2.5 is Runway’s current video model, with 1080p output across text-to-video, image-to-video and video-to-video added in August 2026. Runway now also hosts third-party models, so it increasingly acts as a router rather than a single-model tool. | Gen-3 Alpha Turbo and Gen-4 Aleph were retired on 30 July 2026 and those model IDs now fail; confirm current names before API work. |
| Google Veo High-end generation | Video with audio, references and vertical formats | Veo 3.1 supports text and image inputs, audio-capable output and formats ranging from vertical social clips to higher-resolution production workflows. | Access, resolution and cost vary between Gemini, Flow, Vertex AI and API tiers. |
| Adobe Firefly Creative suite | Brand assets and end-to-end Adobe workflows | Firefly combines Adobe’s own models with partner models, storyboards, editing and brand-oriented production inside a broader creative suite. | Partner models can have different training, rights and credit rules from Adobe models. |
| Kling Creator alternative | Character motion and social video experiments | Kling remains a widely used alternative for text-to-video, image-to-video and creator-focused generation. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Luma Dream Machine Fast ideation | Concept clips and visual iteration | Luma’s Dream Machine and Ray family suit rapid visual exploration and image-to-video workflows. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Pika Social effects | Short-form transformations and creator effects | Pika emphasises approachable effects and short-form creation rather than complex production pipelines. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| HeyGen Avatar localisation | Sales, training and multilingual presenter videos | HeyGen combines avatar presentation, translation and localisation workflows for business content. | Test brand pronunciation, lip sync and disclosure requirements in every target market. |
| Synthesia Enterprise avatars | Governed training and internal communications | Synthesia focuses on controlled presenter-video production, templates, languages and enterprise deployment. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Descript Editor pick | Editing real footage with AI assistance | For many teams, editing captured video with transcript-led tools is more reliable than generating every frame from scratch. | Plans, limits and model availability change frequently; confirm the current vendor page before purchasing. |
| Gemini Omni Emerging multimodal | Conversational video creation from mixed media | Google’s Gemini Omni direction combines images, audio, video and text inputs with conversational generation and editing. | Treat newly released capabilities as emerging until they pass your own production tests. |
Related reading
Sources and verification notes
Primary product documentation checked for this update: