Lumiere is a video diffusion model proposed by researchers from Google, Weizmann Institute of Science, and Tel Aviv University. It aims to generate realistic and stylized videos with the ability to edit them. Users can provide text inputs or upload still images to transform into dynamic videos. The model also supports features like inpainting, cinemagraphs, and stylized generation. Lumiere takes a different approach from existing models by generating the entire temporal duration of the video at once, leading to more realistic and coherent motion. It was trained on a dataset of 30 million videos and is capable of generating 80 frames at 16 fps. Compared to other AI video models, Lumiere produces 5-second videos with higher motion magnitude, temporal consistency, and overall quality. However, it has limitations and cannot generate videos with multiple shots or scene transitions. Lumiere is not yet available for testing, but it shows promise in the rapidly evolving AI video market.
