
| Developer | Zhipu AI |
| Origin | China |
| Released | CogVideoX 1.5 · Nov 2024 |
| Native audio | Add voiceover |
| Max resolution | 1360×768 |
| Max clip length | 6–10s |
| Availability | Open weights · Open weights (2B: Apache 2.0) |
| Available via | Open weights · Z.ai API |
| Pricing from | Free (open) · API |
CogVideoX is best for quick, flexible text- and image-to-video where an accessible, open model is a good fit.
Background. From Zhipu AI and Tsinghua's THUDM, CogVideoX (1.5, November 2024) is one of the most-forked early open video models, runnable on consumer hardware with an ICLR-accepted design.
Open Vivideo and start a new video.
Select CogVideoX as your model — or let the Video Agent pick it for you.
Describe your shot (or upload an image) and set duration and aspect ratio.
Generate with CogVideoX, then refine, add a voice or avatar, and export for any platform.
Each model is one of 30+ in Vivideo — switch per shot to get exactly the look you want.





State-of-the-art motion, realism and native audio.



CogVideoX, from Zhipu AI, is a widely used open-source video model supporting text- and image-to-video. Its openness has made it a community favorite for experimentation.
CogVideoX suits creators who want a flexible, accessible engine for everyday clips and quick iterations across both text and image inputs.
On Vivideo, CogVideoX is one of 30+ models on a single subscription — use it as a flexible everyday option, then escalate to Veo or Sora when a shot needs extra polish, all in one project.
CogVideoX earned its place as one of the most-forked open video models: the 2B variant is Apache 2.0 and commercially free, it runs on as little as ~4–5 GB of VRAM, and the Diffusers, ComfyUI and LoRA ecosystem around it runs deep. CogVideoX text to video and image to video both work well for quick, everyday clips, which made it the default starting point for a generation of tinkerers and researchers.
Be realistic about its age: the open releases date to late 2024, resolution tops out at 1360×768, clips run 6–10 seconds, there's no native audio, and prompts are English-only. Against newer engines it trails on fidelity — WAN and Hunyuan are the stronger open picks today, and Veo 3.1 or Sora 2 win outright on polish. It remains a fine choice for prototyping, B-roll and research.
On Vivideo, a CogVideoX AI video is free to try — daily free generations, no local GPU setup. Use it for fast drafts and layout passes, then re-render the keepers on a heavier model; since output is silent, add an AI voiceover and music in the same project before exporting, or fold clips into a longer 10-minute edit.
You can try CogVideoX free on Vivideo to start — no credit card. Heavier use and premium models are covered by a paid plan.
CogVideoX is an accessible open model that handles both text and image inputs well, making it a flexible everyday option.
Yes — on Vivideo you can use CogVideoX for both text-to-video and image-to-video, and switch to other models per shot.
Yes — Vivideo lets you switch models per shot, so you can mix CogVideoX with other engines in one project.
The 2B open weights are Apache 2.0 — free including commercial use if you self-host. On Vivideo you can also try it free, with daily free generations and no setup.
Open releases top out at 1360×768. For 1080p or 4K delivery, generate the shot with a newer model like Veo 3.1 or LTX-2 instead.