如何获取ModelScope视频生成模型的预测进度并实时展示给用户?
实现视频生成进度实时展示的方案
方法一:重写Pipeline推理逻辑,添加进度回调
ModelScope的文本转视频生成Pipeline底层多基于扩散模型,是逐帧生成视频的。我们可以通过继承原Pipeline类,重写核心推理方法,在每帧生成的节点插入进度回调逻辑。
示例代码:
import torch, random, gc from modelscope.pipelines import pipeline from modelscope.outputs import OutputKeys from modelscope.pipelines.text_to_video_synthesis import TextToVideoSynthesisPipeline from moviepy.editor import VideoFileClip, concatenate_videoclips import datetime # 自定义带进度回调的Pipeline class ProgressTextToVideoPipeline(TextToVideoSynthesisPipeline): def __call__(self, inputs, progress_callback=None, **kwargs): text = inputs['text'] # 根据实际模型生成的视频帧数调整(比如部分模型默认生成16帧) total_frames = 16 current_frame = 0 # 调用原Pipeline的核心推理逻辑 original_output = super().__call__(inputs, **kwargs) # 模拟逐帧生成的进度更新(实际需根据模型内部生成步骤插入回调) for _ in range(total_frames): current_frame += 1 progress = (current_frame / total_frames) * 100 if progress_callback: progress_callback(progress, text) return original_output # 自定义进度更新函数,可替换为前端实时展示逻辑 def update_progress(progress, prompt): print(f"提示词「{prompt}」生成进度: {progress:.1f}%") torch.manual_seed(random.randint(0, 2147483647)) # 初始化自定义Pipeline pipe = ProgressTextToVideoPipeline('text-to-video-synthesis', '/content/drive/MyDrive/BEProject/models') video_clips = [] prompts = ["你的提示词1", "你的提示词2"] # 替换为实际提示词列表 for prompt in prompts: with torch.no_grad(): torch.cuda.empty_cache() gc.collect() # 传入进度回调函数 output_video_path = pipe({'text': prompt}, progress_callback=update_progress)[OutputKeys.OUTPUT_VIDEO] video_clip = VideoFileClip(output_video_path) video_clips.append(video_clip) final = concatenate_videoclips(video_clips) new_video_path = f'/content/videos/{datetime.datetime.now().strftime("%Y-%m-%d_%H:%M:%S")}.mp4' final.write_videofile(new_video_path, codec="libx264")
方法二:基于推理时间估算进度
如果不想修改Pipeline内部逻辑,可以通过预热计算单条提示词的平均生成时间,再实时统计已用时间来估算整体进度,适合对精度要求不高的场景。
示例代码:
import torch, random, gc, time from modelscope.pipelines import pipeline from modelscope.outputs import OutputKeys from moviepy.editor import VideoFileClip, concatenate_videoclips import datetime torch.manual_seed(random.randint(0, 2147483647)) pipe = pipeline('text-to-video-synthesis', '/content/drive/MyDrive/BEProject/models') video_clips = [] prompts = ["你的提示词1", "你的提示词2"] # 预热模型,计算单条提示词的平均生成时间 warmup_prompt = "warmup prompt" start_time = time.time() pipe({'text': warmup_prompt}) avg_time_per_prompt = time.time() - start_time print(f"单条提示词平均生成时间: {avg_time_per_prompt:.2f}秒") total_prompts = len(prompts) for idx, prompt in enumerate(prompts): with torch.no_grad(): torch.cuda.empty_cache() gc.collect() print(f"开始生成第{idx+1}/{total_prompts}条提示词: {prompt}") start_time = time.time() output_video_path = pipe({'text': prompt})[OutputKeys.OUTPUT_VIDEO] elapsed_time = time.time() - start_time # 计算当前整体进度 overall_progress = ((idx + elapsed_time/avg_time_per_prompt) / total_prompts) * 100 print(f"当前整体进度: {overall_progress:.1f}%") video_clip = VideoFileClip(output_video_path) video_clips.append(video_clip) final = concatenate_videoclips(video_clips) new_video_path = f'/content/videos/{datetime.datetime.now().strftime("%Y-%m-%d_%H:%M:%S")}.mp4' final.write_videofile(new_video_path, codec="libx264")
注意事项
- 方法一需要对ModelScope对应Pipeline的内部生成逻辑有基础了解,不同版本的模型生成步骤可能有差异,需对应调整回调插入的位置。
- 方法二的进度估算依赖平均生成时间,若不同提示词的生成耗时差异较大,进度会存在一定误差。
内容的提问来源于stack exchange,提问作者RAMP TSEC
相关产品推荐
相关产品推荐

