Skip to main content
POST
Gemini Native (Video)

Introduction

Veo is Google Vertex AI’s multimodal video generation model. It supports text-to-video (T2V), first-frame constraint, and first-and-last-frame constraint (3.1 series only) for coherent video generation. Use LeapX’s unified video API: submit a task to get a task_id, then query the task to poll status and get the result.

Authentication

string
required
Bearer Token, e.g. Bearer sk-xxxxxxxxxx

Supported models

Call flow

  1. Submit task: POST /v1/video/generations with model, prompt, and Veo-specific parameters.
  2. Poll status: GET /v1/video/generations/{task_id} until status is succeeded or failed.
  3. Get result: On success, url in the response contains the video (Veo may return data:video/mp4;base64,... or an OSS link).

Veo-specific parameters

string
required
Video generation prompt describing the scene and motion.
integer
default:"4"
Video duration in seconds; supported: 4, 6, 8.
string
default:"16:9"
Aspect ratio; only 16:9 and 9:16 supported.
string
default:"1080p"
Resolution: 720p, 1080p.
string
First-frame reference image (URL or Base64) for image-to-video / first-frame constraint.
string
Last-frame reference image (veo-3.1 series only); use with first frame for start-and-end constraint.
boolean
default:"false"
Whether to generate synchronized audio. Fast models ignore this and always include audio.
integer
default:"1"
Number of videos to generate per request, range 1-4.
For more parameters (e.g. personGeneration, addWatermark, seed), see Submit Video Task.

Request examples

Submit Veo task (text-to-video):
Query task status:
For full request/response details and multi-model comparison, see Submit Video Task and Query Video Task.