# Krea.ai Details Krea 2 Image Model Training _Sangwha Lee of Krea.ai details the rigorous data curation and training process behind the Krea 2 image generation model, emphasizing stylistic diversity and efficiency._ **Updated:** 2026-08-22 **Published:** 2026-08-18 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/krea-ai-details-krea-2-image-model-training --- Sangwha Lee from Krea.ai recently shared insights into the training process of their Krea 2 image foundation model, highlighting the critical role of data curation and the pursuit of stylistic diversity. The company has open-sourced a medium version of its model, which has been met with positive reception. Existing ModelsDriver reliable outputs but significant mode collapse, leading to lack of diversityFrom the article 9+ mentionsLee began by contrasting Krea 2 with existing production-grade models like ChatGPT and Nano Banana Pro.drives needKrea 2 GoalEffectaims for faster generation and greater stylistic variation for creative explorationFrom the article 7 mentionsSangwha Lee from Krea.ai recently shared insights into the training process of their Krea 2 image foundation model, highlighting the critical role of data curation and the pursuit of stylistic diversity.requiresData CurationCorerigorous process to ensure stylistic diversity and mitigate 'bad data'From the article 6 mentionsLee emphasized that while architecture is important, "data is quite everything that goes into the model." He stressed that after locking in an architecture, the majority of the effort lies in data curation, ensuring quality and diversity.Diffusion ModelsContextfundamental principles underpin Krea 2's image generation capabilitiesFrom the article 9+ mentionsThe presentation touched upon the fundamental principles of diffusion models, explaining the process of adding noise to an image and training a model to denoise it.Mitigate Bad DataEffectdefining and actively mitigating problematic data for improved model qualityenhancesWorld KnowledgeCoreleveraging advanced training techniques to incorporate broader understandingFrom the article 2 mentionsKrea.ai also incorporated world knowledge into their model by leveraging Wikipedia's PageRank to identify important concepts and ensure their presence in the training dataset.contributes toStylistic DiversityOutcomeachieved through careful data and training, contrasting existing modelsFrom the article 5 mentionsSpecifically for Krea 2, the team focused on maintaining stylistic diversity, even including data like low-resolution CRT videos that might be considered aesthetically poor by some, as they hold value for specific user preferences.leads toOpen-Sourced ModelOutcomeFrom the article 9+ mentionsThe company has open-sourced a medium version of its model, which has been met with positive reception. ## The Quest for Stylistic Diversity Lee began by contrasting Krea 2 with existing production-grade models like ChatGPT and Nano Banana Pro. While these models offer reliable outputs, they often achieve this through significant mode collapse, leading to a lack of diversity. Lee illustrated this with the example of a "burning skull" prompt, where existing models produce consistent but similar results. Krea 2, in contrast, aims for faster generation and greater stylistic variation, catering to users who may not know exactly what they want and need to explore creative possibilities. ## The Foundation of Diffusion Models and Data's Role The presentation touched upon the fundamental principles of diffusion models, explaining the process of adding noise to an image and training a model to denoise it. Lee emphasized that while architecture is important, **"data is quite everything that goes into the model."** He stressed that after locking in an architecture, the majority of the effort lies in data curation, ensuring quality and diversity. Specifically for Krea 2, the team focused on maintaining stylistic diversity, even including data like low-resolution CRT videos that might be considered aesthetically poor by some, as they hold value for specific user preferences. ## Defining and Mitigating "Bad Data" Lee outlined several categories of problematic data for generative models: duplicated samples, over-represented concepts, samples where visual language models fail to capture key aspects, samples that induce biases or artifacts, images with high visual complexity unsuitable for low-resolution training, and AI-generated images. The team employs methods like hash-based deduplication and embedding-based semantic deduplication to address these issues. They also utilize large vision-language models to identify AI-generated content and distill this knowledge into smaller, efficient classifiers. Lee expressed pride in their use of sparse autoencoders (SAEs) for unsupervised tagging, which helps in identifying and filtering undesirable data like watermarks or border artifacts. ## Leveraging World Knowledge and Advanced Training Techniques Krea.ai also incorporated world knowledge into their model by leveraging Wikipedia's PageRank to identify important concepts and ensure their presence in the training dataset. The company's training pipeline is inspired by large language model (LLM) training, progressing through stages like low-to-high resolution pre-training, mid-training, supervised fine-tuning, preference optimization, and reinforcement learning. A key component is the prompt expansion module, which takes user prompts and generates more detailed versions to improve output quality. ## The Future of Image Generation Looking ahead, Lee mentioned Krea.ai's work on multi-expert policy distillation, aiming to merge specialized capabilities from different expert models into a single student model. He also reflected on the evolving architecture of diffusion models, noting a trend towards reversing the traditional encoder-decoder structure. Lee concluded by highlighting the importance of simplicity, scalability, efficiency, and fast iteration in infrastructure and methods, emphasizing the value of drawing inspiration from LLM research. StartupHub.ai data indicates that Krea has a score of 19/100, placing it among competitors like Adulis (20/100) and further behind leaders like Adaptive Biotechnologies (69/100). --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.