Grok Imagine Image to Video: A Reliable Still-to-Clip Workflow
A practical Grok Imagine image-to-video workflow covering still selection, motion prompts, continuity checks, access limits, API polling, failures, and saving the finished clip.
Contents

You may already have a Grok Imagine still you like, yet the video version changes the face, loses a prop, melts the background, or never exposes the same video action shown in someone else’s screenshot. The reliable approach is not endless rerolls: separate the still, motion prompt, continuity review, and account access into four gates.
This guide starts with a button-name-agnostic workflow for the Grok app or web product, then gives a reproducible xAI API path. The outcome is a saved still, a saved short clip, and a repeat loop that tells you whether to fix the source image, the motion instruction, the camera move, or access.
Run the image-to-video workflow once in five steps
For the first pass, use one subject action and one camera action. The goal is to prove the path works before asking for a complex cinematic shot.
- Generate the still. Lock the subject, clothing or key object, composition, background, and lighting; do not pack a long action sequence into this prompt.
- Choose an animation-ready image. Inspect faces, hands, text, edges, and occlusion first. Defects already present in the still usually become more visible in motion.
- Write a separate motion prompt. Describe only what happens next, then list the identity and scene details that must not change.
- Start video from that selected image. Open the image and use the animation or video action that your account actually shows; labels can differ by account, release, and region.
- Download before rerunning. Check motion, subject continuity, camera behavior, and the saved file, then change only one variable for the next version.
Lock identity and composition in the still before adding motion
An animation-ready source image needs to be stable before it needs to be beautiful. Clear facial features, clothing, accessories, held objects, and spatial relationships give you something concrete to preserve and inspect.
Build the prompt in this order: subject and stable traits, composition and camera, background, light and style, then aspect ratio. Leave space in the direction of travel; if you want a push-in, do not start with the face already filling the frame.
Still-image prompt template
Subject and unchanging traits; clothing, accessories, or key object; shot size, camera height, and lens feel; background and lighting; final aspect ratio.
Example
Cinematic waist-up portrait of a young woman in a charcoal coat, a bright yellow bird perched on her shoulder; a busy station crowd forms soft motion blur in the background; eye-level, 50 mm lens feel, shallow depth of field; woman and bird sharp; 16:9.
When choosing among variations, prioritize these checks instead of picking only the most atmospheric frame:
- The face, eyes, hands, and key object are coherent, with separate and readable silhouettes.
- Body parts that need to move are not cropped by the frame or heavily hidden by foreground objects.
- Hair, clothing color, accessories, and subject count are easy to name so you can lock them in the motion prompt.
- The background has no broken text, duplicate people, or dense crossing structures that are likely to drift.
If identity or geometry is already unreliable in the still, regenerate the still first. Do not expect video generation to repair it as a side effect.
Describe change in the motion prompt, not a redesigned scene
A useful motion prompt follows four parts: subject action, camera action, background action, and a lock list. Start with one clear action in each category to reduce the chance of simultaneous face, wardrobe, and scene changes.
Use an action with a visible start and end inside a few seconds. “Turns toward the camera and blinks once” is testable; “shows cinematic emotion” is not. Pick one camera behavior—push in, pull out, pan, or stay still—instead of combining orbit, zoom, and handheld shake.
Motion prompt template
What the subject does over a few seconds; how the camera moves; how the background moves; which face, clothing, object, color, composition, and scene relationships must stay unchanged.
Example
The woman slowly turns toward the camera and blinks once. The yellow bird shifts its feet and opens one wing slightly. The crowd continues moving behind them. The camera slowly pushes in. Keep the woman’s face, charcoal coat, bird color, subject count, framing, and station layout unchanged.
Do not begin with a costume change, weather shift, explosion, multi-person interaction, and complex camera move in one request. Stabilize one verifiable action, then add complexity one item at a time.
Start from the selected image and save both outputs immediately
In the Grok app or web product, the important point is not the exact button label. Confirm that you are continuing from the selected image rather than starting a fresh text-to-video request.
Open the final still, use the animation or image-to-video action currently available to your account, paste the motion prompt, and choose a short, low-complexity first pass. If no action appears, check the app version, signed-in account, plan, region, and any cooldown notice; do not infer access from another person’s screenshot.
As soon as the result finishes, download the still and the clip and keep both prompts in the same record. Hosted media URLs can be temporary, and without local files you cannot compare later versions reliably.
Review five things before you spend another generation
A clip that merely plays is not necessarily usable. Review it in the same order every time so you can distinguish a source-image problem from a motion, camera, or access problem.
- Starting-frame match: The first frame still resembles the chosen image—face, clothing, bird, subject count, and background layout correspond.
- Motion hit: The requested action actually happens with the intended direction, amplitude, and sequence.
- Subject continuity: No face swap, extra or missing limb, color change, disappearing accessory, or merging subjects appears during motion.
- Camera control: The camera performs only the requested move, without an unexpected cut, roll, or violent shake.
- Saved deliverable: The full clip plays, the file is downloaded, and its source still and prompt can be identified.
If identity continuity or file saving fails, do not increase motion yet. Go back to the still or motion prompt; only add complexity after every failure can be traced to a specific gate.
Separate Grok product access, xAI API limits, and user examples
As of 2026-09-28, official API documentation establishes image-to-video capability and the request contract. It does not establish that every Grok web or app account has the same action, quota, or label.
| Evidence | What it can safely tell you |
|---|---|
| Your own Grok app or web screen | Whether this account currently has the action, remaining quota, cooldown notice, and a working download path. |
| Official xAI API documentation | The API model, input formats, asynchronous states, duration, and resolution—not the Grok product’s UI entitlement. |
| One creator’s X post | That the creator produced a short clip at that time, not a universal duration cap, quality level, or free allowance. |
The official image-to-video guide accepts a public image URL, a base64 data URI, or a Files API file_id. The video-generation guide documents a 1–15 second duration range and 480p, 720p, and 1080p output for image-to-video. If aspect_ratio is omitted, the input image ratio is used; overriding it stretches the image. A request returns request_id, moves through pending, done, failed, or expired, and produces a temporary URL when complete.
Treat app quotas, cooldowns, and regional availability as account-specific until your own screen confirms them. API rate limits vary by tier and are not the same thing as generation counts inside the Grok product; before batch use, review current API pricing and the tier shown in your console.
The Frame Zero post on X reports creating a still in Grok Imagine and then a four-second clip from it. That is a useful workflow example, but it is not an official maximum and does not promise the same result for another prompt, account, or source image.
Use the official API when you need a reproducible path
The Bash example below follows the current official contracts for image generation, image-to-video, and asynchronous video polling. It generates and saves a still, starts video generation, captures request_id, polls the terminal states, handles failed/expired, downloads the MP4, and stops after ten minutes.
You need an xAI API account with billing and access enabled, plus curl, jq, and python3. This is Bash. The key is read silently into XAI_API_KEY, so it is not embedded in the script or command history. Keep the first run at duration: 8 and resolution: "720p"; change them only after the full chain succeeds.
set -euo pipefail
command -v curl >/dev/null
command -v jq >/dev/null
command -v python3 >/dev/null
read -rs XAI_API_KEY && export XAI_API_KEY
trap 'unset XAI_API_KEY' EXIT
STILL_PROMPT=$(cat <<'PROMPT'
Cinematic waist-up portrait of a young woman in a charcoal coat, a bright yellow bird perched on her shoulder; a busy station crowd forms soft motion blur in the background; eye-level, 50 mm lens feel, shallow depth of field; woman and bird sharp.
PROMPT
)
MOTION_PROMPT=$(cat <<'PROMPT'
The woman slowly turns toward the camera and blinks once. The yellow bird shifts its feet and opens one wing slightly. The crowd continues moving behind them. The camera slowly pushes in. Keep the woman's face, charcoal coat, bird color, subject count, framing, and station layout unchanged.
PROMPT
)
IMAGE_PAYLOAD=$(jq -n \
--arg prompt "$STILL_PROMPT" \
'{
model: "grok-imagine-image-2.0",
prompt: $prompt,
aspect_ratio: "16:9",
resolution: "1k"
}')
IMAGE_JSON=$(curl --fail-with-body --silent --show-error \
https://api.x.ai/v1/images/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d "$IMAGE_PAYLOAD")
IMAGE_URL=$(jq -r '.data[0].url // empty' <<<"$IMAGE_JSON")
if [ -z "$IMAGE_URL" ]; then
printf '%s\n' "No image URL was found in the image response." >&2
jq . <<<"$IMAGE_JSON" >&2
exit 1
fi
curl --fail-with-body --location --silent --show-error \
"$IMAGE_URL" -o grok-imagine-still.jpg
VIDEO_PAYLOAD=$(jq -n \
--arg prompt "$MOTION_PROMPT" \
--arg image "$IMAGE_URL" \
'{
model: "grok-imagine-video-1.5",
prompt: $prompt,
image: {url: $image},
duration: 8,
resolution: "720p"
}')
START_JSON=$(curl --fail-with-body --silent --show-error \
https://api.x.ai/v1/videos/generations \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $XAI_API_KEY" \
-d "$VIDEO_PAYLOAD")
REQUEST_ID=$(jq -r '.request_id // empty' <<<"$START_JSON")
if [ -z "$REQUEST_ID" ]; then
printf '%s\n' "No request_id was found in the video start response." >&2
jq . <<<"$START_JSON" >&2
exit 1
fi
for _ in $(seq 1 120); do
RESULT=$(curl --fail-with-body --silent --show-error \
"https://api.x.ai/v1/videos/$REQUEST_ID" \
-H "Authorization: Bearer $XAI_API_KEY")
STATUS=$(jq -r '.status // "unknown"' <<<"$RESULT")
case "$STATUS" in
pending)
sleep 5
;;
done)
VIDEO_URL=$(jq -r '.video.url // empty' <<<"$RESULT")
if [ -z "$VIDEO_URL" ]; then
printf '%s\n' "No video URL was found in the completed response." >&2
jq . <<<"$RESULT" >&2
exit 1
fi
curl --fail-with-body --location --silent --show-error \
"$VIDEO_URL" -o grok-imagine-output.mp4
printf '%s\n' "Saved grok-imagine-still.jpg and grok-imagine-output.mp4"
exit 0
;;
failed|expired)
printf 'Video generation ended with status: %s\n' "$STATUS" >&2
jq . <<<"$RESULT" >&2
exit 1
;;
*)
printf 'Unexpected status: %s\n' "$STATUS" >&2
jq . <<<"$RESULT" >&2
exit 1
;;
esac
done
printf '%s\n' "Stopped polling after ten minutes without completion." >&2
exit 1
Success means the script exits with code 0 and the current directory contains both grok-imagine-still.jpg and grok-imagine-output.mp4. HTTP failures stop at curl --fail-with-body; failed or expired prints the complete result and exits instead of waiting forever or silently starting another billed request.
The code is a reproducible template based on the documented request and response fields, not a visual-quality benchmark. It does not guarantee that a stochastic generation will reproduce the sample image. Before batch use, run one request and confirm account access, pricing, rate limits, output quality, and download behavior.
Fix the most likely failure before changing every parameter
Change one upstream variable per attempt so the next result can actually identify the cause.
There is no image-to-video action
First make sure you opened the final selected image rather than a new text conversation. Then check the app version, signed-in account, plan, region, and any on-screen cooldown notice. Do not copy a button name from someone else’s screenshot; users with API access can use the official image-to-video endpoint instead.
The face changes or a key object disappears
Return to a source image with clearer silhouettes and less occlusion, reduce the action amplitude, and end the prompt with an explicit lock list for face, clothing, color, subject count, and the key object. If frame one already differs, adding more prose is rarely the right fix.
The clip is nearly static
Replace abstract language with a visible transition that has a start and an end, such as “turns from looking left to facing the camera and blinks once.” Keep one subject action and one camera action so requests do not cancel one another.
The background melts or the camera runs away
Remove background motion first and keep only the subject action; set the camera to still or a single slow push-in. Crowds, text, fences, and crossing lines are drift-prone, so simplify the source background when necessary.
The API stays pending, then fails or expires
Poll at a fixed interval and do not submit duplicates just because several minutes have passed. After failed or expired, inspect the response, key and account state, and whether the input URL is still reachable, then create a new request. Stop at your timeout instead of polling indefinitely.
The completed URL later stops working
Official docs describe generated image and video URLs as temporary. Download as soon as the request reaches done, name the local files by version, and do not use the hosted URL as a permanent asset store.
Hold the still constant and change one variable for version two
A useful rerun answers one question; “try again” answers none.
- Keep the same source still and identity lock list, and change only motion amplitude to see whether continuity improves.
- After motion is stable, change the camera—for example, from locked-off to a slow push-in—without replacing the source image.
- After the camera is stable, adjust duration or resolution and record processing time, terminal status, and download outcome.
- Store version number, still filename, motion prompt, parameters, and review result on one line so you can roll back.
This isolates whether the problem came from the still, the motion instruction, the camera, or access. Stabilize one simple chain before adding a second action or a more complicated scene.
References and scope
The pages below were checked on 2026-09-28. Official pages support API capabilities and fields; the X post is used only as an individual workflow report.
- xAI Image Generation — model,
/v1/images/generations, temporary image URLs, and configuration. - xAI Image-to-Video — accepted image inputs and a
grok-imagine-video-1.5example. - xAI Video Generation — asynchronous flow, states, duration, resolution, and temporary output URLs.
- xAI Videos REST API Reference —
request_idand the result endpoint. - xAI Release Notes — the 2026-07-31 record for
grok-imagine-video-1.5modalities. - Frame Zero post on X — a personal workflow report, not an official product limit or general quality claim.