CogVideoX
Open text-to-video and image-to-video generation models with inference and fine-tuning code.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About CogVideoX
CogVideoX is a family of video generation models, released with the earlier CogVideo, that create short clips from a text prompt or a starting image. The repository includes inference and fine-tuning code, plus tools for working with the models.
The 5B model can be tried online through Hugging Face and ModelScope Spaces, or downloaded and run on your own GPU. Code and model weights are under separate licences. It suits developers and researchers experimenting with local video generation.
Strengths
- Text-to-video and image-to-video generation
- Fine-tuning code included
- Online demos available before installing
Limitations
- Needs a powerful GPU to run locally
- Model weights have their own licence terms
- Research code rather than a finished app
Details
- Pricing
- FreeFree to download; model use is governed by a separate model licence.
- License
- Proprietary
- Developer
- Zhipu AI (zai-org)
- Platforms
- Linux, Command line
- How it runs
- Downloadable app
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Developers and researchers running video generation models locally
- Categories
- Local AI tools, AI video
- Last verified
- Added
- Sources
Alternatives to CogVideoX
Compare allSoftware that can replace CogVideoX for an important use case, and what changes if you switch.
HunyuanVideo
Tencent's large text-to-video generation model, with weights and inference code you run yourself.
HunyuanVideo is Tencent's text-to-video model with published weights and a Gradio interface, but needs a high-end GPU and uses a custom licence.
Wan 2.2
Open, advanced large-scale video generation model for local use.
Wan 2.2 is Apache-2.0 licensed with several model variants to fit different GPUs and runs on Windows too, though it lacks an official end-user GUI.
Open-Sora
An open-source video generation project releasing models, training code and tools for text-to-video.
Open-Sora publishes models and training code with a Gradio interface and a large community, but requires powerful GPUs and machine learning experience.
LTX-2
Open local video model that generates synchronized audio and video.
LTX-2 generates synchronized audio and video together, includes an official LoRA trainer and runs on Windows, though its license terms are not standard.
Mochi 1 (Genmo)
Open state-of-the-art video generation model from research lab Genmo.
Mochi 1 has Apache-2.0 open weights with consumer-GPU support through ComfyUI, though its repository shows no activity since late 2025.
FramePack
Local image-to-video generator that predicts frames progressively so it runs on consumer GPUs.
FramePack is Apache-2.0 image-to-video code that keeps workload constant for longer videos on consumer GPUs, but lacks text-to-video and fine-tuning code.
CogVideoX as an alternative
Listings that name CogVideoX as an alternative.
Wan2GP
A local AI video generator built to run open video and image models on low-memory GPUs.
CogVideoX is research code for text-to-video and image-to-video with fine-tuning support, needing a powerful GPU rather than targeting low-memory cards.
Similar software
Related functionality, not necessarily a direct replacement.
LivePortrait
An open research project that animates a still portrait using the motion from a driving video.
Google Veo
Google DeepMind's video generation model that creates clips with native audio from text prompts.
Kling AI
An AI studio for generating images and videos from text, images and references.
Runway
A web creative platform for generating and editing video, images and audio with AI.