You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python实现音乐播放中实时调整tempo?解决pydub无法实时干预问题

实时调整音频播放速度的解决方案

问题分析

你当前的pydub代码通过speedup()预先完成全音频的变速处理,再用阻塞式的play()播放,这种预处理模式无法在播放过程中动态修改速度参数,完全满足不了“实时调整tempo并即时生效”的需求。

实现思路

要实现实时变速,必须切换为流式分段处理+实时播放的模式:

  • 将音频切割成极小的片段(比如100ms级)
  • 对即将播放的片段实时执行变速处理
  • 播放处理后的片段,同时允许外部修改速度参数
  • 后续片段自动沿用最新的速度设置

依赖安装

先安装所需依赖库:

pip install pydub sounddevice numpy

代码实现示例

以下代码结合pydub做音频变速(保持音高),用sounddevice实现实时流式播放,同时通过线程监听键盘输入动态调整速度:

import sounddevice as sd
from pydub import AudioSegment
import numpy as np
import threading
import sys

# 全局可修改的速度参数,默认1.0倍速
current_speed = 1.0
# 播放终止标记
stop_playback = False
# 当前播放到原始音频的位置
current_pos = 0

def audio_callback(outdata, frames, time, status):
    global current_speed, current_pos, audio_data, stop_playback
    
    if status:
        print(status, file=sys.stderr)
    
    # 根据当前速度计算需要读取的原始音频帧数
    original_frames = int(frames / current_speed)
    # 检测是否播放到音频末尾
    if current_pos + original_frames >= len(audio_data):
        outdata[:] = 0
        stop_playback = True
        return
    
    # 提取当前要处理的音频片段
    segment = audio_data[current_pos:current_pos+original_frames]
    current_pos += original_frames
    
    # 将numpy数组转为pydub的AudioSegment,执行变速(保持音高)
    seg = AudioSegment(
        segment.tobytes(),
        frame_rate=sample_rate,
        sample_width=audio_data.dtype.itemsize,
        channels=channels
    )
    sped_seg = seg.speedup(playback_speed=current_speed)
    
    # 转换回numpy数组并填充到输出缓冲区
    sped_array = np.frombuffer(sped_seg.raw_data, dtype=audio_data.dtype)
    # 确保长度匹配输出缓冲区,不足则补0
    if len(sped_array) < frames * channels:
        sped_array = np.pad(sped_array, (0, frames * channels - len(sped_array)), mode='constant')
    outdata[:] = sped_array.reshape(-1, channels)

def speed_control_thread():
    global current_speed, stop_playback
    while not stop_playback:
        cmd = input("\n输入新速度(如0.5/1.0/2.0),输入q停止播放:")
        if cmd.lower() == 'q':
            stop_playback = True
            break
        try:
            new_speed = float(cmd)
            if new_speed > 0:
                current_speed = new_speed
                print(f"已切换至{current_speed}倍速")
        except ValueError:
            print("无效输入,请输入正数或q")

# 加载目标音频文件
sound = AudioSegment.from_wav('mymusic.wav')
sample_rate = sound.frame_rate
channels = sound.channels
audio_data = np.frombuffer(sound.raw_data, dtype=np.int16)

# 配置sounddevice的输出参数
sd.default.samplerate = sample_rate
sd.default.channels = channels

# 启动速度控制线程(后台运行)
threading.Thread(target=speed_control_thread, daemon=True).start()

# 启动实时流式播放
print("开始播放,可随时输入新速度调整")
with sd.Stream(callback=audio_callback):
    while not stop_playback:
        pass
print("播放结束")

关键说明

  • 流式处理:通过sounddevice.Stream的回调函数,每次仅处理即将播放的小片段音频,避免提前预处理整个文件
  • 动态速度生效:全局变量current_speed可被控制线程实时修改,后续所有音频片段都会立即使用新速度
  • 音高保持:使用pydub.speedup()实现变速时不改变音高,避免单纯重采样导致的音调偏移
  • 非阻塞控制:单独的线程监听用户输入,不会阻塞播放流程

轻量替代方案

如果不需要保持音高,可直接用numpy重采样实现更轻量化的实时变速(会改变音高),只需替换回调中的变速逻辑为:

# 直接重采样实现变速(改变音高)
sped_array = np.interp(
    np.linspace(0, len(segment), frames),
    np.arange(len(segment)),
    segment
).astype(audio_data.dtype)

内容的提问来源于stack exchange,提问作者Maanuv Jagtiani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:15:44