You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中通过变速调整.wav文件时长匹配目标时长

问题背景

我正在为生成的字幕文件生成对应的TTS音频,先对字幕文件进行清洗处理(时间戳行以##开头),代码如下:

duration = VideoFileClip(f"media/{name}.mp4").duration
for i in range(len(lines)):
    if '##' in lines[i]:
        timestamp = lines[i].replace('## ', '').replace(' :', '').replace(' :\n', '')
        start_time, end_time = map(float, timestamp.split(" - "))
        if lines[i + 1] == '\n':
            subtitle_entries.append({"start": start_time, "end": end_time,
                                     "text": "<empty>"})
    elif lines[i] == '\n':
        continue
    else:
        subtitle_entries.append({"start": start_time, "end": end_time,
                                 "text": lines[i].lstrip() if lines[i].lstrip() != '' else "<empty>"})
        data += lines[i][:-1] + ' '

final_entries = []
if subtitle_entries[0]['start'] != 0:
    final_entries.append({"start": 0, "end": subtitle_entries[0]['start'], "text": "<empty>"})

for i in range(len(subtitle_entries)):
    if i != 0 and subtitle_entries[i - 1]['end'] != subtitle_entries[i]['start']:
        final_entries.append({"start": subtitle_entries[i - 1]['end'], "end": subtitle_entries[i]['start'],
                              "text": "<empty>"})
    final_entries.append(subtitle_entries[i])

if subtitle_entries[-1]['end'] != total_dur:
    final_entries.append({"start": subtitle_entries[-1]['end'], "end": total_dur, "text": "<empty>"})

处理后生成字幕条目传入TTS模型,生成音频的逻辑如下:

from pydub import AudioSegment
for entry in final_entries:
    if entry['text'] == "<empty>":
        audio_files.append(AudioSegment.silent(duration=(entry['end'] - entry['start']) * 1000))
    else:
        tts.tts_to_file(text=entry['text'], file_path=f"media/{name}/{entry['end']}.wav")

        audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav")
        original_dur = len(audio_file)
        expected_dur = (entry['end'] - entry['start']) * 1000

        # Speedup or slow down the audio file here

        audio_files.append(audio_file)
遇到的问题

需要将每个TTS生成的.wav文件变速,使其匹配字幕条目的目标时长,实现音画同步。尝试使用librosa时出现报错,报错代码及日志如下:

报错代码

for entry in final_entries:
    if entry['text'] == "<empty>":
        audio_files.append(AudioSegment.silent(duration=(entry['end'] - entry['start']) * 1000))
    else:
        tts.tts_to_file(text=entry['text'], file_path=f"media/{name}/{entry['end']}.wav")

        audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav")
        original_durr = len(audio_file)
        expected_durr = (entry['end'] - entry['start']) * 1000

        y, sr = librosa.load(f"media/{name}/{entry['end']}.wav")
        y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr))
        sf.write(f"media/{name}/{entry['end']}.wav", y_output, sr)

        audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav")
        audio_files.append(audio_file)

错误日志

Internal Server Error: /video_downloader/download/
Traceback (most recent call last):
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\exception.py", line 55, in inner
    response = get_response(request)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\base.py", line 197, in _get_response
    response = wrapped_callback(request, *callback_args, **callback_kwargs)
  File "C:\Academics\transcriber\video_downloader\views.py", line 158, in download_video
    audio_file = tts_convertor(translated_file, target_language, name)
  File "C:\Academics\transcriber\video_processor\processor.py", line 148, in tts_convertor
    y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr))
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\effects.py", line 245, in time_stretch
    stft_stretch = core.phase_vocoder(
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\core\spectrum.py", line 1457, in phase_vocoder
    d_stretch[..., t] = util.phasor(phase_acc, mag=mag)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\util\utils.py", line 2602, in phasor
    z = _phasor_angles(angles)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\numba\np\ufunc\dufunc.py", line 190, in __call__
    return super().__call__(*args, **kws)
numpy.core._exceptions._UFuncNoLoopError: ufunc '_phasor_angles' did not contain a loop with signature matching types <class 'numpy.dtype[float32]'> -> None
ERROR:django.request:Internal Server Error: /video_downloader/download/
Traceback (most recent call last):
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\exception.py", line 55, in inner
    response = get_response(request)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\base.py", line 197, in _get_response
    response = wrapped_callback(request, *callback_args, **callback_kwargs)
  File "C:\Academics\transcriber\video_downloader\views.py", line 158, in download_video
    audio_file = tts_convertor(translated_file, target_language, name)
  File "C:\Academics\transcriber\video_processor\processor.py", line 148, in tts_convertor
    y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr))
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\effects.py", line 245, in time_stretch
    stft_stretch = core.phase_vocoder(
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\core\spectrum.py", line 1457, in phase_vocoder
    d_stretch[..., t] = util.phasor(phase_acc, mag=mag)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\util\utils.py", line 2602, in phasor
    z = _phasor_angles(angles)
  File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\numba\np\ufunc\dufunc.py", line 190, in __call__
    return super().__call__(*args, **kws)
numpy.core._exceptions._UFuncNoLoopError: ufunc '_phasor_angles' did not contain a loop with signature matching types <class 'numpy.dtype[float32]'> -> None
[03/Dec/2023 14:29:23] "POST /video_downloader/download/ HTTP/1.1" 500 141577
解决方案

1. 修复Librosa的报错

Librosa报错是因为音频数据类型不兼容,加载时指定dtype=np.float64即可解决,同时注意修正变速方向:

import numpy as np
import librosa
import soundfile as sf

# ... 其他代码 ...

y, sr = librosa.load(f"media/{name}/{entry['end']}.wav", dtype=np.float64)
# time_stretch的rate参数为原时长/目标时长:rate>1加速,rate<1减速
rate = original_durr / expected_durr
y_output = librosa.effects.time_stretch(y, rate=rate)
sf.write(f"media/{name}/{entry['end']}.wav", y_output, sr)

2. 使用Pydub实现变速

Pydub可直接对AudioSegment对象变速,无需反复读写文件,更高效:

from pydub import AudioSegment
from pydub.effects import speedup

# ... 其他代码 ...

audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav")
original_dur = len(audio_file)
expected_dur = (entry['end'] - entry['start']) * 1000

speed_factor = original_dur / expected_dur
if speed_factor != 1.0:
    # 加速场景直接用speedup
    if speed_factor > 1:
        audio_file = speedup(audio_file, playback_speed=speed_factor)
    # 减速场景需安装pydub-effects扩展:pip install pydub-effects
    else:
        from pydub_effects import change_speed
        audio_file = change_speed(audio_file, speed_factor)
    
    # 修正微小时长误差
    adjusted_dur = len(audio_file)
    if adjusted_dur < expected_dur:
        audio_file += AudioSegment.silent(duration=expected_dur - adjusted_dur)
    elif adjusted_dur > expected_dur:
        audio_file = audio_file[:expected_dur]

audio_files.append(audio_file)

3. 使用MoviePy实现变速

MoviePy适合和视频处理流程整合,自带音调保持的变速功能:

from moviepy.editor import AudioFileClip, vfx

# ... 其他代码 ...

audio_clip = AudioFileClip(f"media/{name}/{entry['end']}.wav")
original_dur = audio_clip.duration
expected_dur = entry['end'] - entry['start']

speed_factor = original_dur / expected_dur
# 变速并保持音调
audio_clip = audio_clip.fx(vfx.speedx, speed_factor)
# 写入调整后的音频文件
audio_clip.write_audiofile(f"media/{name}/{entry['end']}_adjusted.wav")
# 转换为pydub对象继续处理
audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}_adjusted.wav")

audio_files.append(audio_file)

内容的提问来源于stack exchange,提问作者Leofierus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 21:25:52