Stability AI has released an early preview of its Stable Diffusion 3.0, a next-generation text-to-image generative AI model. The company has been consistently improving its image models over the past year, and the new model aims to provide better image quality and performance. It also focuses on improving typography, an area where previous models have struggled. Stable Diffusion 3.0 is based on a new architecture called diffusion transformers, which enable a new era of image generation. The model is being developed in multiple sizes, ranging from 800M to 8B parameters. Stability AI has also been experimenting with other approaches, such as the Würstchen architecture in Stable Cascade. The improved typography in Stable Diffusion 3.0 is achieved through the use of transformer architecture and additional text encoders. The model is initially demonstrated as a text-to-image technology but will serve as the foundation for future visual models, including video and 3D image generation.
