You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于时间戳列表裁剪音频文件并解决索引类型错误问题

我有一个时长很长的音频文件,已将文本文件中手动标注的关键片段秒级起止时间转换为嵌套列表[[start1,end1],[start2,end2],...]。需要遍历该列表,从原始音频中依次裁剪对应时间段的片段,且需保留浮点型时间戳,后续将提取MFCC等音频特征。尝试使用以下scipy代码时出现“索引必须为整数”的错误,现咨询如何将秒级时长关联到音频数组的索引:

fs1, y1 = scipy.io.wavfile.read(file_path)
l1 = numpy.array(annotation_list)   
newWavFileAsList = []
for elem in l1:
    startRead = elem[0]
    endRead = elem[1]
    newWavFileAsList.extend(y1[startRead:endRead])
newWavFile = numpy.array(newWavFileAsList)

scipy.io.wavfile.write(sample, fs1, newWavFile)
解决方案

为啥报错?因为音频数组y1的索引得是整数,但你直接用了浮点型的秒级时间戳。要把秒数转成数组索引,核心靠采样率(fs1)——采样率就是每秒采集的音频样本数,所以:

  • 起始索引 = 起始秒数 × 采样率
  • 结束索引 = 结束秒数 × 采样率

计算完记得转成整数,数组不认浮点数索引。至于是用四舍五入还是直接取整,看你需求:要精准对齐用numpy.round(),想直接向下取整就用int()强制转换。

改好的代码如下:

import scipy.io.wavfile
import numpy as np

fs1, y1 = scipy.io.wavfile.read(file_path)
annotation_list = [[start1, end1], [start2, end2], ...]  # 你的标注列表
newWavFileAsList = []

for start_sec, end_sec in annotation_list:
    # 秒转数组索引,转成整数
    start_idx = int(np.round(start_sec * fs1))
    end_idx = int(np.round(end_sec * fs1))
    # 防止索引超出音频数组范围
    start_idx = max(0, start_idx)
    end_idx = min(len(y1), end_idx)
    # 裁剪片段并加入列表
    newWavFileAsList.extend(y1[start_idx:end_idx])

newWavFile = np.array(newWavFileAsList)
scipy.io.wavfile.write(sample, fs1, newWavFile)

额外提两点:

  • 如果是多声道音频(比如立体声),y1的形状是(样本数, 声道数),上面的代码照样能用,切片是按样本维度处理的。
  • 要保留浮点时间戳很简单,直接存你原始标注的start_sec和end_sec就行,后续提取MFCC时直接关联这些值就行,不用从裁剪后的音频反推。

内容的提问来源于stack exchange,提问作者Medium

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 05:52:56