CogVideoX

Open text-to-video and image-to-video generation models with inference and fine-tuning code.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About CogVideoX

CogVideoX is a family of video generation models, released with the earlier CogVideo, that create short clips from a text prompt or a starting image. The repository includes inference and fine-tuning code, plus tools for working with the models.

The 5B model can be tried online through Hugging Face and ModelScope Spaces, or downloaded and run on your own GPU. Code and model weights are under separate licences. It suits developers and researchers experimenting with local video generation.

Strengths

  • Text-to-video and image-to-video generation
  • Fine-tuning code included
  • Online demos available before installing

Limitations

  • Needs a powerful GPU to run locally
  • Model weights have their own licence terms
  • Research code rather than a finished app

Details

Pricing
FreeFree to download; model use is governed by a separate model licence.
License
Proprietary
Developer
Zhipu AI (zai-org)
Platforms
Linux, Command line
How it runs
Downloadable app
Account
Not required
Works offline
Yes
Best suited for
Developers and researchers running video generation models locally
Last verified
Added

Alternatives to CogVideoX

Compare all

Software that can replace CogVideoX for an important use case, and what changes if you switch.

  • HunyuanVideo

    Tencent's large text-to-video generation model, with weights and inference code you run yourself.

    HunyuanVideo is Tencent's text-to-video model with published weights and a Gradio interface, but needs a high-end GPU and uses a custom licence.

  • Wan 2.2

    Open, advanced large-scale video generation model for local use.

    Wan 2.2 is Apache-2.0 licensed with several model variants to fit different GPUs and runs on Windows too, though it lacks an official end-user GUI.

  • Open-Sora

    An open-source video generation project releasing models, training code and tools for text-to-video.

    Open-Sora publishes models and training code with a Gradio interface and a large community, but requires powerful GPUs and machine learning experience.

  • LTX-2

    Open local video model that generates synchronized audio and video.

    LTX-2 generates synchronized audio and video together, includes an official LoRA trainer and runs on Windows, though its license terms are not standard.

  • Mochi 1 (Genmo)

    Open state-of-the-art video generation model from research lab Genmo.

    Mochi 1 has Apache-2.0 open weights with consumer-GPU support through ComfyUI, though its repository shows no activity since late 2025.

  • FramePack

    Local image-to-video generator that predicts frames progressively so it runs on consumer GPUs.

    FramePack is Apache-2.0 image-to-video code that keeps workload constant for longer videos on consumer GPUs, but lacks text-to-video and fine-tuning code.

CogVideoX as an alternative

Listings that name CogVideoX as an alternative.

  • Wan2GP

    A local AI video generator built to run open video and image models on low-memory GPUs.

    FreeProprietaryWindowsLinux

    CogVideoX is research code for text-to-video and image-to-video with fine-tuning support, needing a powerful GPU rather than targeting low-memory cards.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing