Python中通过变速调整.wav文件时长匹配目标时长
问题背景
我正在为生成的字幕文件生成对应的TTS音频,先对字幕文件进行清洗处理(时间戳行以##开头),代码如下:
duration = VideoFileClip(f"media/{name}.mp4").duration for i in range(len(lines)): if '##' in lines[i]: timestamp = lines[i].replace('## ', '').replace(' :', '').replace(' :\n', '') start_time, end_time = map(float, timestamp.split(" - ")) if lines[i + 1] == '\n': subtitle_entries.append({"start": start_time, "end": end_time, "text": "<empty>"}) elif lines[i] == '\n': continue else: subtitle_entries.append({"start": start_time, "end": end_time, "text": lines[i].lstrip() if lines[i].lstrip() != '' else "<empty>"}) data += lines[i][:-1] + ' ' final_entries = [] if subtitle_entries[0]['start'] != 0: final_entries.append({"start": 0, "end": subtitle_entries[0]['start'], "text": "<empty>"}) for i in range(len(subtitle_entries)): if i != 0 and subtitle_entries[i - 1]['end'] != subtitle_entries[i]['start']: final_entries.append({"start": subtitle_entries[i - 1]['end'], "end": subtitle_entries[i]['start'], "text": "<empty>"}) final_entries.append(subtitle_entries[i]) if subtitle_entries[-1]['end'] != total_dur: final_entries.append({"start": subtitle_entries[-1]['end'], "end": total_dur, "text": "<empty>"})
处理后生成字幕条目传入TTS模型,生成音频的逻辑如下:
from pydub import AudioSegment for entry in final_entries: if entry['text'] == "<empty>": audio_files.append(AudioSegment.silent(duration=(entry['end'] - entry['start']) * 1000)) else: tts.tts_to_file(text=entry['text'], file_path=f"media/{name}/{entry['end']}.wav") audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav") original_dur = len(audio_file) expected_dur = (entry['end'] - entry['start']) * 1000 # Speedup or slow down the audio file here audio_files.append(audio_file)
遇到的问题
需要将每个TTS生成的.wav文件变速,使其匹配字幕条目的目标时长,实现音画同步。尝试使用librosa时出现报错,报错代码及日志如下:
报错代码
for entry in final_entries: if entry['text'] == "<empty>": audio_files.append(AudioSegment.silent(duration=(entry['end'] - entry['start']) * 1000)) else: tts.tts_to_file(text=entry['text'], file_path=f"media/{name}/{entry['end']}.wav") audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav") original_durr = len(audio_file) expected_durr = (entry['end'] - entry['start']) * 1000 y, sr = librosa.load(f"media/{name}/{entry['end']}.wav") y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr)) sf.write(f"media/{name}/{entry['end']}.wav", y_output, sr) audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav") audio_files.append(audio_file)
错误日志
Internal Server Error: /video_downloader/download/ Traceback (most recent call last): File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\exception.py", line 55, in inner response = get_response(request) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\base.py", line 197, in _get_response response = wrapped_callback(request, *callback_args, **callback_kwargs) File "C:\Academics\transcriber\video_downloader\views.py", line 158, in download_video audio_file = tts_convertor(translated_file, target_language, name) File "C:\Academics\transcriber\video_processor\processor.py", line 148, in tts_convertor y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr)) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\effects.py", line 245, in time_stretch stft_stretch = core.phase_vocoder( File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\core\spectrum.py", line 1457, in phase_vocoder d_stretch[..., t] = util.phasor(phase_acc, mag=mag) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\util\utils.py", line 2602, in phasor z = _phasor_angles(angles) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\numba\np\ufunc\dufunc.py", line 190, in __call__ return super().__call__(*args, **kws) numpy.core._exceptions._UFuncNoLoopError: ufunc '_phasor_angles' did not contain a loop with signature matching types <class 'numpy.dtype[float32]'> -> None ERROR:django.request:Internal Server Error: /video_downloader/download/ Traceback (most recent call last): File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\exception.py", line 55, in inner response = get_response(request) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\django\core\handlers\base.py", line 197, in _get_response response = wrapped_callback(request, *callback_args, **callback_kwargs) File "C:\Academics\transcriber\video_downloader\views.py", line 158, in download_video audio_file = tts_convertor(translated_file, target_language, name) File "C:\Academics\transcriber\video_processor\processor.py", line 148, in tts_convertor y_output = librosa.effects.time_stretch(y, rate=(expected_durr / original_durr)) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\effects.py", line 245, in time_stretch stft_stretch = core.phase_vocoder( File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\core\spectrum.py", line 1457, in phase_vocoder d_stretch[..., t] = util.phasor(phase_acc, mag=mag) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\librosa\util\utils.py", line 2602, in phasor z = _phasor_angles(angles) File "C:\Users\leofi\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.10_qbz5n2kfra8p0\LocalCache\local-packages\Python310\site-packages\numba\np\ufunc\dufunc.py", line 190, in __call__ return super().__call__(*args, **kws) numpy.core._exceptions._UFuncNoLoopError: ufunc '_phasor_angles' did not contain a loop with signature matching types <class 'numpy.dtype[float32]'> -> None [03/Dec/2023 14:29:23] "POST /video_downloader/download/ HTTP/1.1" 500 141577
解决方案
1. 修复Librosa的报错
Librosa报错是因为音频数据类型不兼容,加载时指定dtype=np.float64即可解决,同时注意修正变速方向:
import numpy as np import librosa import soundfile as sf # ... 其他代码 ... y, sr = librosa.load(f"media/{name}/{entry['end']}.wav", dtype=np.float64) # time_stretch的rate参数为原时长/目标时长:rate>1加速,rate<1减速 rate = original_durr / expected_durr y_output = librosa.effects.time_stretch(y, rate=rate) sf.write(f"media/{name}/{entry['end']}.wav", y_output, sr)
2. 使用Pydub实现变速
Pydub可直接对AudioSegment对象变速,无需反复读写文件,更高效:
from pydub import AudioSegment from pydub.effects import speedup # ... 其他代码 ... audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}.wav") original_dur = len(audio_file) expected_dur = (entry['end'] - entry['start']) * 1000 speed_factor = original_dur / expected_dur if speed_factor != 1.0: # 加速场景直接用speedup if speed_factor > 1: audio_file = speedup(audio_file, playback_speed=speed_factor) # 减速场景需安装pydub-effects扩展:pip install pydub-effects else: from pydub_effects import change_speed audio_file = change_speed(audio_file, speed_factor) # 修正微小时长误差 adjusted_dur = len(audio_file) if adjusted_dur < expected_dur: audio_file += AudioSegment.silent(duration=expected_dur - adjusted_dur) elif adjusted_dur > expected_dur: audio_file = audio_file[:expected_dur] audio_files.append(audio_file)
3. 使用MoviePy实现变速
MoviePy适合和视频处理流程整合,自带音调保持的变速功能:
from moviepy.editor import AudioFileClip, vfx # ... 其他代码 ... audio_clip = AudioFileClip(f"media/{name}/{entry['end']}.wav") original_dur = audio_clip.duration expected_dur = entry['end'] - entry['start'] speed_factor = original_dur / expected_dur # 变速并保持音调 audio_clip = audio_clip.fx(vfx.speedx, speed_factor) # 写入调整后的音频文件 audio_clip.write_audiofile(f"media/{name}/{entry['end']}_adjusted.wav") # 转换为pydub对象继续处理 audio_file = AudioSegment.from_wav(f"media/{name}/{entry['end']}_adjusted.wav") audio_files.append(audio_file)
内容的提问来源于stack exchange,提问作者Leofierus
相关产品推荐
相关产品推荐

