Vidomaker AI

720p with sound · No watermark · From $0.62 a video

One sentence. One video.

Describe the scene, get an HD video in two minutes.

Loading…

107/1000

« A red fox crosses a snowy forest at sunrise, the camera follows it from the side, snowflakes gently falling »

Creating a video from a sentence: what changed in 2026

Two years ago, AI video generation gave strange results: six-fingered hands, faces melting from one frame to the next, impossible camera moves. The 2026 models crossed a threshold nobody expected so soon. Characters keep their appearance from shot to shot, light behaves as it does in the real world, water, smoke and fabric move the way they should, and sound is generated together with the image: footsteps, wind, engines, city ambience.

This generator relies on these latest-generation models, in particular ByteDance's Seedance family, chosen for the fidelity of faces and the precision of movement. You write a sentence in your own language, we translate and enrich it for the model, and a five- or ten-second 720p video with sound is ready in one to two minutes.

This guide explains how a video model works, how to write a description that gives the result you have in mind, what the formats and styles offer, and why we chose credits rather than a subscription.

How a model turns text into video

A video generation model doesn't “film” anything: it learns, from millions of described videos, the relationship between words and sequences of images. When you give it a description, it starts from random noise and progressively denoises it, frame by frame and above all through time, until it converges on a sequence consistent with the text. It's the same principle as image generation, with one more dimension: continuity between successive frames.

That temporal continuity is what separates a good model from a bad one. A weak model produces beautiful but independent frames, and the video shakes or changes subject. A recent model reasons about the whole sequence at once: the fox crossing the forest stays the same fox, the shadow follows the light, the camera glides without jolts. That's why we only use models from the last quarter, and why we'll replace them as soon as a better one comes out.

Sound follows the same logic: the model generates an audio track consistent with what it shows. A wave produces a rush, a street at night a distant hum of traffic, a fireplace its crackle. It isn't library music added afterwards, it's an ambience computed for the scene, which gives the clips a realism the image alone doesn't have.

Writing a description that works

The quality of your video depends first on your description. The model is very good at inventing the details you don't specify, but it can't guess what you have in mind. A sentence like “a cat” gives any cat in any setting; “a ginger cat asleep on a window ledge, curtain moving in the wind, golden late-afternoon light” gives exactly that scene.

Think like a director describing a shot to their crew. There are four things to cover: the subject and what it does, the setting, the light or time of day, and the camera movement. You don't have to give them all, but each one reduces randomness. Camera movement is the one most often forgotten, and it's the one that changes the result the most: “the camera slowly moves forward”, “drone shot descending”, “static close-up”, “lateral tracking shot”.

Avoid the keyword lists of 2023 image generators: “8k, ultra detailed, masterpiece” adds nothing for a video model and takes the place of real information. Also avoid asking for two contradictory actions in five seconds, or a scene that changes location: a short clip tells a single moment. If your story has three moments, make three videos.

Finally, write naturally, in your own language. Our server translates your description into English, the language in which the models are most precise, and adds the style indications you chose. You have nothing more to do, and you can reread the prompt used under each video to understand what worked.

  • One subject, one action: “a woman runs on a beach” rather than “a woman runs, then swims, then goes home”
  • Setting and time: “a Lisbon alley, early morning, ground still wet”
  • Camera movement: “the camera slowly circles the subject”
  • Mood: “light mist, soft colors, silence”
  • What to avoid: no on-screen text, no logo, no named real person

Formats, durations and styles: what to choose

The format depends on where the video will be watched. 16:9 landscape is the natural format for YouTube, a website or a presentation. 9:16 portrait is the format of TikTok, Instagram Reels and Shorts: if your video is meant for a phone, choose it from the start rather than cropping afterwards, the model composes the scene for the requested format. 1:1 square suits classic Instagram posts, animated avatars and thumbnails.

The standard duration is five seconds, which matches a film shot or a social media clip: long enough for an action to unfold, short enough to stay sharp. The ten-second version costs two credits and suits scenes where something needs to build up, like a sunrise or a camera gradually revealing a place. For an edit, five seconds per shot is the right rhythm.

Styles are indications added to your description. “Realistic” aims for the look of a cinema camera, “Cinematic” adds a more contrasted image and slow movements, “Anime” applies the line work and colors of Japanese animation, “3D” gives an animated-film render, “Drone” frames a wide aerial shot, “Vintage” adds the grain and colors of old film stock. You can choose none: the model then interprets your description freely.

The “High quality” mode of the selector is different from the “Cinematic” style: it's a more advanced model, our best, sold in dedicated packs and counting as two standard videos. It adds finesse on faces, hands and complex movements. For a landscape, an object or an atmosphere, the standard model is more than enough; keep high quality for scenes with characters in close-up.

What you can do with your videos

The videos you generate are yours and are delivered without a watermark, as MP4 files that play everywhere. You can publish them on social media, embed them in a presentation or a website, use them in an edit or an ad. The most common uses among our users: animated backgrounds for TikTok or YouTube videos, product visuals for an online store, animated illustrations for a course or training, mood shots for a short film, and personalised gifts.

Every video has a share page with a short link, a preview that displays properly on WhatsApp, Messenger, X or Discord, and the prompt used so that the people who receive it can create their own. This page isn't indexed by search engines: only the people you send the link to can see it.

Files are kept on our servers for thirty days, long enough to download and use them, then deleted: we don't build a library of your creations. Download what matters. A deleted video can't be regenerated identically, every generation is unique.

Why credits, and not a subscription

A video genuinely costs money to produce: a few tens of cents of computing on specialised graphics cards, a hundred times more than a chat message. An “unlimited” subscription at that price would be a lie, like the ones sold by services that slow down the queue as soon as you really use it. We prefer a simple, honest system: one credit, one video. The price per video is written on every pack, with no abstract units and no conversion table.

Credits don't subscribe you to anything. You buy a pack when you need it, use it at your own pace for twelve months, and nothing is charged afterwards. The first pack, ten videos for $9.99, is an afternoon of creation; the thirty-video pack is the favourite of those who post every week; the eighty-video pack is for projects and professionals. Within seven days of purchase, unused credits are refundable on simple request.

Finally, no sign-up is required before buying: you pay, your credits are available immediately on this device, and an account is created with your payment email so you can find them elsewhere. Nobody should have to register just to look around.

The limits, stated plainly

Video generation in 2026 is impressive, not perfect. On-screen text is often illegible: if you need a title, add it afterwards in an editing program. Hands remain the weak point of every model in close-up scenes. Very complex scenes, with many characters interacting, can lose coherence. And the model interprets: two generations of the same description will give two different videos, which is also what makes it fun.

Some requests are refused before any generation, without using a credit: sexual content, explicit violence, identifiable real people, minors in inappropriate situations, incitement to hatred. These are the rules of the models we use, and ours. When a description is refused, the message tells you clearly and you can rephrase it.

If a generation fails for a technical reason, the credit is refunded automatically and immediately; you have nothing to ask for. The announced generation time is a real median of the latest videos produced, not a marketing promise: it varies with server load, and we'd rather tell you “about 90 seconds” and keep to it than “instant” and disappoint.

Tips to go further

Work in variations. A first generation shows you how the model understands your description; the second, with one or two words changed, is almost always better. Your description stays in the field after each video, so iterating is immediate. The most effective creators produce three versions of a shot and keep the best, rather than looking for the perfect description on the first try.

For an edit, generate all your shots in the same format and the same style, describing the same light: the result will have a visual unity viewers feel without being able to explain it. Alternate shot sizes as in cinema: a wide shot to set the scene, a medium shot for the action, a close-up for emotion.

And look at the examples at the top of this page with their exact description: they are real videos produced by this generator, with no cosmetic selection. Clicking one copies its description into the field; change the subject, keep the structure, and you have a solid starting point.

Frequently asked questions about the video generator

How much does a video cost?

One credit per 5-second standard video. The first pack, 10 videos for $9.99, works out at $1.00 a video; the Studio pack goes down to $0.62 a video. A 10-second or high-quality video counts as two standard videos; High quality packs are expressed directly in HQ videos.

Is there a free trial?

No, and we'd rather say so plainly: every video has a real computing cost, and free trials are paid for elsewhere with endless queues or watermarks. Instead, the examples at the top of this page are real videos made by this generator, with their exact description, and unused credits are refundable within 7 days.

Do I need to create an account?

Not before paying. Right after your purchase you choose a password: your account is created with the email used for payment, with your credits, so you can find your credits and videos on all your devices. A link sent by email also lets you do it later.

How long does a video take?

Usually one to two minutes for 5 seconds in 720p. The time shown during generation is the real median of the latest videos produced. You can leave the page: the video waits for you in “My videos”, and your browser can notify you.

Can I use the videos commercially?

Yes. Videos are delivered as MP4 without a watermark and you can publish them, edit them or use them in an ad. The only limits: no identifiable real people and no content forbidden by our rules, which is refused before generation.

What happens if a generation fails?

Your credit is refunded automatically, no need to ask. A description refused by the safety rules is never charged: the refusal happens before any computing.

Do credits expire?

They stay valid for twelve months after purchase, with no renewal and no recurring charge. Generated videos are kept on our servers for thirty days: download the ones you want to keep.