Text to Video Guide: Prompts, Models, and Settings

Direct answer: Text-to-video starts with a prompt and no source image. It is the right workflow when you want the model to invent the opening composition, or when you are still exploring the subject, camera angle, environment, and look.

Fact-checked September 10, 2026. This page covers the no-image workflow. If a useful first frame already exists, use the image-to-video guide. For category selection, read the AI video generator guide. For cost math, use AI video generator pricing.

Text-to-video workflow

StepDecisionPractical default
Define the shotWhat happens in one clip?One subject and one main action
Write the promptWhat should appear and move?Subject + action + setting + camera + light
Pick a modelWhich capability boundary matters?Compare duration, audio, draft, and price
Pick settingsWhere will the video be used?Five seconds, 720p, and the delivery ratio
Check costIs this attempt worth the quote?Review the complete quote before generating
IterateWhat single change would improve it?Change one prompt or setting variable

Write a shot, not a script

Use this structure:

Subject and action, in setting, camera movement or framing, lighting and visual treatment, motion constraint.

Example:

A silver road bicycle leaning outside a quiet bakery at sunrise, its front wheel turning gently in the breeze, slow lateral camera slide, soft amber light and realistic reflections, restrained natural motion.

That prompt has one subject, one small action, one environment, one camera instruction, and one treatment. It does not ask the same five-second clip to introduce a location, transform the bicycle, add a crowd, and cut to a close-up.

Prompt variants by job

Product detail: Matte-black headphones on a stone pedestal, subtle rotation while a narrow highlight travels across the ear cups, locked medium close-up, dark studio background, precise premium lighting.

Atmosphere: Rain falling across a neon-lit side street at night, one cyclist crossing the frame, fixed wide camera, reflections moving across wet pavement, restrained cinematic realism.

Vertical social clip: A chef placing the final herb garnish on a plated dish, gentle handheld push-in, bright natural kitchen light, crisp appetizing detail, one continuous action.

Choose settings deliberately

magicdoor.ai exposes a text prompt, duration in five-second increments, 16:9/9:16/1:1, and 720p/1080p across Wan 3, Kling 3.0, and FLUX 3.

SettingChoose based onCost implication
DurationHow long the single action needsPrice scales by output second
RatioFinal placementChoose 16:9, 9:16, or 1:1
ResolutionExploration versus delivery1080p costs more than 720p
Generated audioWhether the shot needs model-generated soundKling 3.0 and FLUX 3 support it; Kling audio costs more
FLUX draftWhether this is an exploratory 720p attempt$0.06/second instead of normal 720p at $0.17/second

Wan 3 supports up to 30 seconds at $0.05/second for 720p or $0.10/second for 1080p. Kling 3.0 supports up to 15 seconds with distinct silent/audio rates. FLUX 3 supports up to 20 seconds, optional generated audio, and a 720p draft mode. This is a settings comparison, not a quality ranking.

A low-waste first pass

  • Wan 3: $0.25 at 720p and $0.50 at 1080p for five seconds.
  • Kling 3.0: $0.84 silent or $1.26 with audio at 720p; $1.12 silent or $1.68 with audio at 1080p.
  • FLUX 3: $0.30 in 720p draft, $0.85 at normal 720p, and $1.45 at 1080p. Audio does not currently change its price.

Once the composition and motion work, decide whether delivery genuinely needs more duration or 1080p. A larger setting does not repair a vague prompt.

Troubleshooting text-to-video

SymptomNext attempt
The scene feels chaoticKeep one subject and one action
Motion changes directionUse one camera move and one subject movement
Framing is unpredictableAdd shot size, camera position, and subject placement
The result is staticUse a concrete verb and name what moves
Cost climbs quicklyReturn to five seconds and 720p; change one variable
Audio is missingUse Kling 3.0 or FLUX 3 and enable generated audio

When a composition becomes valuable, save the best frame or prepare an image and move to the image-to-video workflow.

Before opening the creator

Only active or trialing subscribers can generate video. The $6/month base includes $1 in credit, and usage comes from included credit or account balance. magicdoor.ai shows the full quote before each run. Finished videos remain private until deleted, and up to three jobs can reserve funds concurrently.

Anonymous visitors to /videos are redirected to sign in. Have a prompt ready, then continue through the account and subscription flow.

Create a video

FAQ

What is text-to-video?

Text-to-video creates a clip from a written description without a source image. The prompt defines the subject, action, setting, camera, and visual treatment, while model settings define duration, shape, resolution, and available audio behavior.

What should a text-to-video prompt include?

Include one subject, one main action, the setting, one camera instruction, and the light or visual treatment. Add constraints only when they clarify the shot, and avoid packing multiple scenes into one short generation.

How should I test a text-to-video idea cheaply?

Start with a five-second 720p run, review the quoted price before generation, and change one variable per retry. FLUX 3 also offers a 720p draft setting; every retry is another charged run.

Sources

Accessed September 10, 2026.

  • magicdoor.ai pricing for subscription and included-credit details.
  • The current magicdoor.ai video model registry and quote logic for the prompt-first input, exposed settings, model limits, and customer prices.
  • Wan 3, Kling 3.0, and FLUX 3 official model pages as supporting workflow documentation. magicdoor.ai copy is limited to controls exposed by the product.

Copyright © 2026 magicdoor.ai