glow-shadow
Tegas Logo

Text-to-video & photo-to-video

Generate videos up to 30 seconds — from text or a single photo

Tegas creates a complete clip in one generation: a story that develops, camera motion, native scene audio and speaking characters. No editing, no stock footage, no prompt engineering — the built-in AI assistant writes the script for you.

What Tegas clip generation can do

30 seconds

in a single generation — steps of 5, 10, 15, 20, 25, 30 seconds

1080p

three quality tiers: 480p drafts, 720p and 1080p for publishing

Native sound

ambience, engines, footsteps and lip-synced character speech — born with the video

9:16 · 16:9 · 1:1

vertical, horizontal and square formats for every platform

Text to video

Describe the scene in plain words — the Wan 3.0 model generates one continuous clip with a beginning, development and finale. For long clips the Tessa assistant writes the story phase by phase with timecodes.

AI script assistant

Tell your idea in your own words — Tessa asks a couple of questions, builds a professional prompt with a mini-story and fills the form for you: description, music, voice-over and duration.

A story, not a loop

A 15–30 second clip is structured in phases: each changes the action, camera or light. You get a story with a finale instead of a repeating fragment.

Speaking characters

The model generates character speech with lip sync, plus music and natural scene sounds — audio is created together with the picture.

Photo to video

One photo becomes a living scene: the frame comes to life, the camera starts moving, new elements enter the shot. Works with product shots, interiors, landscapes and old family photos.

First & last frame

Upload two frames — the model generates a smooth transition between them. Perfect for before/after videos: renovations, makeovers, packaging.

The assistant sees your photo

Tessa analyzes the uploaded image and builds the story from what is actually in the frame: location, subjects, light and mood.

Ready-made templates

“Product Ad”, “Delicious Food”, “Living Memory”, “Epic Trailer” and more — one click instead of a prompt. iPhone HEIC photos convert automatically.

Reference photos: characters and products stay consistent

Attach up to 10 photos to your text prompt — a character from different angles, a product or a style sample. Wan 3.0 keeps them consistent across every frame.

  • The same hero in every scene — for episodic clips and brand mascots
  • Accurate product appearance in ads — no distorted packaging
  • A consistent visual style from a sample image

How to create a video: 3 steps

1

Describe your idea or upload a photo

In plain words, no prompts — or press “Help me with the prompt” and Tessa builds everything for you.

2

Pick duration and quality

5 to 30 seconds, 480p/720p/1080p, and the right format for your platform: Reels, Shorts, TikTok or web.

3

Get a finished video with sound

Generation takes a few minutes. The clip is saved to your history — download and publish.

AI video generation FAQ

Up to 30 seconds in a single continuous generation — with steps of 5, 10, 15, 20, 25 and 30 seconds. It is one coherent scene with a storyline, not a stitch of short fragments.

Yes. Upload a photo and the model brings it to life with motion, camera work and scene sound. JPEG, PNG, WebP and iPhone HEIC are supported (HEIC converts automatically).

Yes — audio is generated together with the picture: natural scene sounds, music, and character speech with lip sync. Fully silent clips are also possible.

You provide two images — the opening and the final frame — and the model generates a smooth video transition between them. Popular for before/after videos.

Attach up to 10 photos of a character, product or style to your prompt — the model keeps their appearance consistent in every frame, so heroes and products do not morph between scenes.

No. The built-in Tessa assistant turns an idea told in plain words into a professional script with phases, camera work and sound, and fills the generation form automatically.

Pay with tokens: from 2 tokens per second in 480p to 8 tokens per second in 1080p. A five-second draft starts at 10 tokens. Tokens come with subscription plans or top-up packages.

Tegas clips run on Wan 3.0 — a state-of-the-art video model with native audio, up to 30-second single-pass generation, first/last frame control and reference images.

Try it for free

Sign up in a minute — with an email code or Google. Build your first clip with the assistant in 5 minutes.

Start generating
AI Video Generator from Text & Photo — 30-second Clips with Sound | Tegas · Tegas AI