Creating a video from a sentence: what changed in 2026
Two years ago, AI video generation gave strange results: six-fingered hands, faces melting from one frame to the next, impossible camera moves. The 2026 models crossed a threshold nobody expected so soon. Characters keep their appearance from shot to shot, light behaves as it does in the real world, water, smoke and fabric move the way they should, and sound is generated together with the image: footsteps, wind, engines, city ambience.
This generator relies on five of these latest-generation models: Seedance 1.0 Fast and Seedance 2.0 Fast from ByteDance, Grok Imagine, Wan 3.0 from Alibaba and Kling 2.6, which animates an image. You write a sentence in your own language; it is translated and enriched with what a video model needs (camera, light, motion), then sent to the model you chose. A few minutes later, you download an MP4 in 720p, with sound when the model makes it, without a watermark.
This guide explains how a video model works, how to write a description that gives the result you have in mind, what the formats and styles offer, and how the credits of the monthly plans work.
How a model turns text into video
A video generation model doesn't “film” anything: it learns, from millions of described videos, the relationship between words and sequences of images. When you give it a description, it starts from random noise and progressively denoises it, frame by frame and above all through time, until it converges on a sequence consistent with the text. It's the same principle as image generation, with one more dimension: continuity between successive frames.
That temporal continuity is what separates a good model from a bad one. A weak model produces beautiful but independent frames, and the video shakes or changes subject. A recent model reasons about the whole sequence at once: the fox crossing the forest stays the same fox, the shadow follows the light, the camera glides without jolts. That's why we only use models from the last quarter, and why we'll replace them as soon as a better one comes out.
Sound follows the same logic: the model generates an audio track consistent with what it shows. A wave produces a rush, a street at night a distant hum of traffic, a fireplace its crackle. It isn't library music added afterwards, it's an ambience computed for the scene, which gives the clips a realism the image alone doesn't have.
Writing a description that works
The quality of your video depends first on your description. The model is very good at inventing the details you don't specify, but it can't guess what you have in mind. A sentence like “a cat” gives any cat in any setting; “a ginger cat asleep on a window ledge, curtain moving in the wind, golden late-afternoon light” gives exactly that scene.
Think like a director describing a shot to their crew. There are four things to cover: the subject and what it does, the setting, the light or time of day, and the camera movement. You don't have to give them all, but each one reduces randomness. Camera movement is the one most often forgotten, and it's the one that changes the result the most: “the camera slowly moves forward”, “drone shot descending”, “static close-up”, “lateral tracking shot”.
Avoid the keyword lists of 2023 image generators: “8k, ultra detailed, masterpiece” adds nothing for a video model and takes the place of real information. Also avoid asking for two contradictory actions in five seconds, or a scene that changes location: a short clip tells a single moment. If your story has three moments, make three videos.
Finally, write naturally, in your own language. Our server translates your description into English, the language in which the models are most precise, and adds the style indications you chose. You have nothing more to do, and you can reread the prompt used under each video to understand what worked.
- One subject, one action: “a woman runs on a beach” rather than “a woman runs, then swims, then goes home”
- Setting and time: “a Lisbon alley, early morning, ground still wet”
- Camera movement: “the camera slowly circles the subject”
- Mood: “light mist, soft colors, silence”
- What to avoid: no on-screen text, no logo, no named real person
Formats, durations and styles: what to choose
The format depends on where the video will be watched. 16:9 landscape is the natural format for YouTube, a website or a presentation. 9:16 portrait is the format of TikTok, Instagram Reels and Shorts: if your video is meant for a phone, choose it from the start rather than cropping afterwards, the model composes the scene for the requested format. 1:1 square suits classic Instagram posts, animated avatars and thumbnails.
The standard duration is five seconds, which matches a film shot or a social media clip: long enough for an action to unfold, short enough to stay sharp. The ten-second version uses twice as many credits and suits scenes where something needs to build up, like a sunrise or a camera gradually revealing a place. For an edit, five seconds per shot is the right rhythm.
Styles are indications added to your description. “Realistic” aims for the look of a cinema camera, “Cinematic” adds a more contrasted image and slow movements, “Anime” applies the line work and colors of Japanese animation, “3D” gives an animated-film render, “Drone” frames a wide aerial shot, “Vintage” adds the grain and colors of old film stock. You can choose none: the model then interprets your description freely.
The model matters as much as the style. Seedance 1.0 Fast is the quickest and cheapest, silent, perfect for trying an idea. Grok Imagine brings lively, expressive scenes with sound. Wan 3.0 follows long descriptions closely. Seedance 2.0 Fast is our most realistic model, for faces and natural motion. Kling 2.6 starts from your own image, a photo or a drawing, and brings it to life: the video takes the format of the image.
What you can do with your videos
The videos you generate are yours and are delivered without a watermark, as MP4 files that play everywhere. You can publish them on social media, embed them in a presentation or a website, use them in an edit or an ad. The most common uses among our users: animated backgrounds for TikTok or YouTube videos, product visuals for an online store, animated illustrations for a course or training, mood shots for a short film, and personalised gifts.
Every video has a share page with a short link, a preview that displays properly on WhatsApp, Messenger, X or Discord, and the prompt used so that the people who receive it can create their own. This page isn't indexed by search engines: only the people you send the link to can see it.
Files are kept on our servers for thirty days, long enough to download and use them, then deleted: we don't build a library of your creations. Download what matters. A deleted video can't be regenerated identically, every generation is unique.
Why credits, in a monthly plan
A video genuinely costs money to produce: a few tens of cents of computing on specialised graphics cards, a hundred times more than a chat message. An “unlimited” plan at a small price would be a lie, like the ones sold by services that slow down the queue as soon as you really use it. We prefer a simple, honest system: credits that match the real cost of each video, shown on every option before you click.
Each plan includes a number of credits every month: the first, $9.99 a month, gives 400 credits; the 1,000-credit plan is the favourite of those who post every week; the 2,000-credit plan is for projects and professionals. A 5-second video uses 11 to 70 credits depending on the model, shown before you click, and ten seconds twice as many. Credits are renewed with each payment; those you don't use expire at the end of the month.
No commitment: you change plan or cancel in one click from “My account”, the cancellation takes effect at the end of the paid month, and nothing more is charged. You can look around, watch the examples and read everything without an account; the free account only comes in when you create or subscribe.
The limits, stated plainly
Video generation in 2026 is impressive, not perfect. On-screen text is often illegible: if you need a title, add it afterwards in an editing program. Hands remain the weak point of every model in close-up scenes. Very complex scenes, with many characters interacting, can lose coherence. And the model interprets: two generations of the same description will give two different videos, which is also what makes it fun.
Some requests are refused before any generation, without using a credit: sexual content, explicit violence, identifiable real people, minors in inappropriate situations, incitement to hatred. These are the rules of the models we use, and ours. When a description is refused, the message tells you clearly and you can rephrase it.
If a generation fails for a technical reason, the credit is refunded automatically and immediately; you have nothing to ask for. The announced generation time is a real median of the latest videos produced, not a marketing promise: it varies with server load, and we'd rather tell you “about 90 seconds” and keep to it than “instant” and disappoint.
Tips to go further
Work in variations. A first generation shows you how the model understands your description; the second, with one or two words changed, is almost always better. Your description stays in the field after each video, so iterating is immediate. The most effective creators produce three versions of a shot and keep the best, rather than looking for the perfect description on the first try.
For an edit, generate all your shots in the same format and the same style, describing the same light: the result will have a visual unity viewers feel without being able to explain it. Alternate shot sizes as in cinema: a wide shot to set the scene, a medium shot for the action, a close-up for emotion.
And look at the examples at the top of this page with their exact description: they are real videos produced by this generator, with no cosmetic selection. Clicking one copies its description into the field; change the subject, keep the structure, and you have a solid starting point.