You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取ModelScope视频生成模型的预测进度并实时展示给用户?

实现视频生成进度实时展示的方案

方法一:重写Pipeline推理逻辑,添加进度回调

ModelScope的文本转视频生成Pipeline底层多基于扩散模型,是逐帧生成视频的。我们可以通过继承原Pipeline类,重写核心推理方法,在每帧生成的节点插入进度回调逻辑。

示例代码:

import torch, random, gc
from modelscope.pipelines import pipeline
from modelscope.outputs import OutputKeys
from modelscope.pipelines.text_to_video_synthesis import TextToVideoSynthesisPipeline
from moviepy.editor import VideoFileClip, concatenate_videoclips
import datetime

# 自定义带进度回调的Pipeline
class ProgressTextToVideoPipeline(TextToVideoSynthesisPipeline):
    def __call__(self, inputs, progress_callback=None, **kwargs):
        text = inputs['text']
        # 根据实际模型生成的视频帧数调整(比如部分模型默认生成16帧)
        total_frames = 16
        current_frame = 0

        # 调用原Pipeline的核心推理逻辑
        original_output = super().__call__(inputs, **kwargs)
        
        # 模拟逐帧生成的进度更新(实际需根据模型内部生成步骤插入回调)
        for _ in range(total_frames):
            current_frame += 1
            progress = (current_frame / total_frames) * 100
            if progress_callback:
                progress_callback(progress, text)
        
        return original_output

# 自定义进度更新函数,可替换为前端实时展示逻辑
def update_progress(progress, prompt):
    print(f"提示词「{prompt}」生成进度: {progress:.1f}%")

torch.manual_seed(random.randint(0, 2147483647))
# 初始化自定义Pipeline
pipe = ProgressTextToVideoPipeline('text-to-video-synthesis', '/content/drive/MyDrive/BEProject/models')

video_clips = []
prompts = ["你的提示词1", "你的提示词2"]  # 替换为实际提示词列表

for prompt in prompts:
    with torch.no_grad(): 
        torch.cuda.empty_cache()
    gc.collect()
    # 传入进度回调函数
    output_video_path = pipe({'text': prompt}, progress_callback=update_progress)[OutputKeys.OUTPUT_VIDEO]
    video_clip = VideoFileClip(output_video_path)
    video_clips.append(video_clip)

final = concatenate_videoclips(video_clips)
new_video_path = f'/content/videos/{datetime.datetime.now().strftime("%Y-%m-%d_%H:%M:%S")}.mp4'
final.write_videofile(new_video_path, codec="libx264")

方法二:基于推理时间估算进度

如果不想修改Pipeline内部逻辑,可以通过预热计算单条提示词的平均生成时间,再实时统计已用时间来估算整体进度,适合对精度要求不高的场景。

示例代码:

import torch, random, gc, time
from modelscope.pipelines import pipeline
from modelscope.outputs import OutputKeys
from moviepy.editor import VideoFileClip, concatenate_videoclips
import datetime

torch.manual_seed(random.randint(0, 2147483647))
pipe = pipeline('text-to-video-synthesis', '/content/drive/MyDrive/BEProject/models')

video_clips = []
prompts = ["你的提示词1", "你的提示词2"]

# 预热模型,计算单条提示词的平均生成时间
warmup_prompt = "warmup prompt"
start_time = time.time()
pipe({'text': warmup_prompt})
avg_time_per_prompt = time.time() - start_time
print(f"单条提示词平均生成时间: {avg_time_per_prompt:.2f}秒")

total_prompts = len(prompts)
for idx, prompt in enumerate(prompts):
    with torch.no_grad(): 
        torch.cuda.empty_cache()
    gc.collect()
    
    print(f"开始生成第{idx+1}/{total_prompts}条提示词: {prompt}")
    start_time = time.time()
    output_video_path = pipe({'text': prompt})[OutputKeys.OUTPUT_VIDEO]
    elapsed_time = time.time() - start_time
    
    # 计算当前整体进度
    overall_progress = ((idx + elapsed_time/avg_time_per_prompt) / total_prompts) * 100
    print(f"当前整体进度: {overall_progress:.1f}%")
    
    video_clip = VideoFileClip(output_video_path)
    video_clips.append(video_clip)

final = concatenate_videoclips(video_clips)
new_video_path = f'/content/videos/{datetime.datetime.now().strftime("%Y-%m-%d_%H:%M:%S")}.mp4'
final.write_videofile(new_video_path, codec="libx264")

注意事项

  • 方法一需要对ModelScope对应Pipeline的内部生成逻辑有基础了解,不同版本的模型生成步骤可能有差异,需对应调整回调插入的位置。
  • 方法二的进度估算依赖平均生成时间,若不同提示词的生成耗时差异较大,进度会存在一定误差。

内容的提问来源于stack exchange,提问作者RAMP TSEC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 20:57:35