SDXL指定图像尺寸后触发requires_aesthetics_score报错求助
Stable Diffusion XL 指定自定义尺寸后报错的解决方法
问题情况
首次使用Stable Diffusion,按官方教程操作原本正常,但指定图像height和width后出现报错。尝试在refiner加载到cuda前添加requires_aesthetics_score=True仍无法解决,报错信息如下:
ValueError Traceback (most recent call last) Cell In[74], line 1 ----> 1 refiner_image = refiner( 2 prompt="cartoon of colorful monsters frolocking in a dark spooky graveyard with tombstones and graves behind a castle", 3 num_inference_steps=n_steps, 4 denoising_end=high_noise_frac, 5 image=img 6 ).images[0] File c:\Users\Mark\anaconda3\envs\auto_content_creator\lib\site-packages\torch\utils\_contextlib.py:115, in context_decorator..decorate_context(*args, **kwargs) 112 @functools.wraps(func) 113 def decorate_context(*args, **kwargs): 114 with ctx_factory(): --> 115 return func(*args, **kwargs) File c:\Users\Mark\anaconda3\envs\auto_content_creator\lib\site-packages\diffusers\pipelines\stable_diffusion_xl\pipeline_stable_diffusion_xl_img2img.py:910, in StableDiffusionXLImg2ImgPipeline.__call__(self, prompt, prompt_2, image, strength, num_inference_steps, denoising_start, denoising_end, guidance_scale, negative_prompt, negative_prompt_2, num_images_per_prompt, eta, generator, latents, prompt_embeds, negative_prompt_embeds, pooled_prompt_embeds, negative_pooled_prompt_embeds, output_type, return_dict, callback, callback_steps, cross_attention_kwargs, guidance_rescale, original_size, crops_coords_top_left, target_size, aesthetic_score, negative_aesthetic_score) 908 # 8. Prepare added time ids & embeddings 909 add_text_embeds = pooled_prompt_embeds --> 910 add_time_ids, add_neg_time_ids = self._get_add_time_ids( 911 original_size, 912 crops_coords_top_left, 913 target_size, 914 aesthetic_score, 915 negative_aesthetic_score, 916 dtype=prompt_embeds.dtype, 917 ) 918 add_time_ids = add_time_ids.repeat(batch_size * num_images_per_prompt, 1) 920 if do_classifier_free_guidance: File c:\Users\Mark\anaconda3\envs\auto_content_creator\lib\site-packages\diffusers\pipelines\stable_diffusion_xl\pipeline_stable_diffusion_xl_img2img.py:613, in StableDiffusionXLImg2ImgPipeline._get_add_time_ids(self, original_size, crops_coords_top_left, target_size, aesthetic_score, negative_aesthetic_score, dtype) 607 expected_add_embed_dim = self.unet.add_embedding.linear_1.in_features 609 if ( 610 expected_add_embed_dim > passed_add_embed_dim 611 and (expected_add_embed_dim - passed_add_embed_dim) == self.unet.config.addition_time_embed_dim 612 ): --> 613 raise ValueError( 614 f"Model expects an added time embedding vector of length {expected_add_embed_dim}, but a vector of {passed_add_embed_dim} was created. Please make sure to enable `requires_aesthetics_score` with `pipe.register_to_config(requires_aesthetics_score=True)` to make sure `aesthetic_score` {aesthetic_score} and `negative_aesthetic_score` {negative_aesthetic_score} is correctly used by the model." 615 ) 616 elif ( 617 expected_add_embed_dim < passed_add_embed_dim 618 and (passed_add_embed_dim - expected_add_embed_dim) == self.unet.config.addition_time_embed_dim 619 ): 620 raise ValueError( 621 f"Model expects an added time embedding vector of length {expected_add_embed_dim}, but a vector of {passed_add_embed_dim} was created. Please make sure to disable `requires_aesthetics_score` with `pipe.register_to_config(requires_aesthetics_score=False)` to make sure `target_size` {target_size} is correctly used by the model." 622 ) ValueError: Model expects an added time embedding vector of length 2816, but a vector of 2560 was created. Please make sure to enable `requires_aesthetics_score` with `pipe.register_to_config(requires_aesthetics_score=True)` to make sure `aesthetic_score` 6.0 and `negative_aesthetic_score` 2.5 is correctly used by the model.
用户原代码:
from diffusers import StableDiffusionXLPipeline, DiffusionPipeline import torch import os base = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float16, variant="fp16", use_safetensors=True ) base.to("cuda") refiner = DiffusionPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-refiner-1.0", **base.components ) refiner.to("cuda") refiner.register_to_config(requires_aesthetics_score=True) n_steps = 40 high_noise_frac = 0.8 def text_to_image(prompt): base_image = base( prompt=prompt, num_inference_steps=n_steps, denoising_end=high_noise_frac, output_type="latent", height=640, width=1536 ).images refiner_image = refiner( prompt=prompt, num_inference_steps=n_steps, denoising_end=high_noise_frac, image=base_image ).images[0] return refiner_image img = text_to_image("cartoon of colorful monsters frolocking in a dark spooky graveyard with tombstones and graves behind a castle")
解决方法
问题根源是配置时机错误和缺少尺寸参数传递,修改如下:
- 调整
register_to_config的位置,必须在refiner移到CUDA之前设置 - 调用refiner时,显式传入
original_size和target_size参数,匹配你指定的height和width
修改后的完整代码:
from diffusers import StableDiffusionXLPipeline, DiffusionPipeline import torch import os base = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", torch_dtype=torch.float16, variant="fp16", use_safetensors=True ) base.to("cuda") refiner = DiffusionPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-refiner-1.0", **base.components ) # 移到cuda前先配置 refiner.register_to_config(requires_aesthetics_score=True) refiner.to("cuda") n_steps = 40 high_noise_frac = 0.8 # 定义尺寸变量,避免重复写 img_height = 640 img_width = 1536 def text_to_image(prompt): base_image = base( prompt=prompt, num_inference_steps=n_steps, denoising_end=high_noise_frac, output_type="latent", height=img_height, width=img_width ).images refiner_image = refiner( prompt=prompt, num_inference_steps=n_steps, denoising_end=high_noise_frac, image=base_image, # 添加尺寸参数 original_size=(img_height, img_width), target_size=(img_height, img_width) ).images[0] return refiner_image img = text_to_image("cartoon of colorful monsters frolocking in a dark spooky graveyard with tombstones and graves behind a castle")
说明
requires_aesthetics_score=True需要在模型移到CUDA前配置,否则配置不会生效- 传入
original_size和target_size是为了让refiner正确生成对应尺寸的时间嵌入向量,避免维度不匹配的报错 - 把尺寸定义成变量可以减少重复代码,方便后续修改
内容的提问来源于stack exchange,提问作者Mark Dabler
相关产品推荐
相关产品推荐

