You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

流式加载HuggingFace数据集时音频文件无法下载加载求助

流式加载Common Voice数据集时音频加载失败的解决方案

问题描述

流式加载HuggingFace数据集mozilla-foundation/common_voice_17_0的代码如下:

language = "en"
buffer_size = 100
streaming_dataset = load_dataset("mozilla-foundation/common_voice_17_0", language, split="train",
                                 use_auth_token=True, trust_remote_code=True, streaming=True)

dataset = streaming_dataset.shuffle(
    seed=1137, buffer_size=buffer_size
)
common_voice = buffer_dataset(dataset, buffer_size)

sample = random.choice(common_voice)

audio_file = sample["path"]

所有步骤执行正常,但调用torchaudio.load(audio_file)加载音频时出现报错:

Traceback (most recent call last):
  ...
  "... .venv/lib/python3.10/site-packages/soundfile.py", line 1216, in _open
    raise LibsndfileError(err, prefix="Error opening {0!r}: ".format(self.name))
soundfile.LibsndfileError: Error opening 'en_train_14/common_voice_en_18837621.mp3': System error.

非流式加载小数据集分片可正常运行,但'en'这类大数据集无法采用该方案,推测是streaming=True时音频文件未完成下载导致的问题。


解决方案

1. 直接使用数据集自带的音频字段加载

流式模式下,样本的audio字段已包含音频的原始数据和采样率,无需依赖本地文件路径加载:

# 从sample的audio字段直接提取波形和采样率
waveform = torch.tensor(sample["audio"]["array"])
sample_rate = sample["audio"]["sampling_rate"]

这种方式会自动流式获取音频数据,无需等待完整文件下载,完全适配流式场景。

2. 手动下载指定音频文件(适用于必须用本地文件的场景)

如果需要将音频文件保存到本地后再加载,可以利用样本隐含的远程文件信息,通过数据集构建器单独下载:

from datasets import load_dataset_builder

# 获取数据集构建器,用于处理文件下载
ds_builder = load_dataset_builder("mozilla-foundation/common_voice_17_0", language)
# 获取当前样本对应的远程音频文件URL
remote_file_url = sample["_remote_files"][0]["url"]
# 指定本地保存路径
local_file_path = "./target_audio.mp3"

# 下载文件到本地
ds_builder.download(remote_file_url, local_file_path)

# 现在可以正常加载本地文件
waveform, sample_rate = torchaudio.load(local_file_path)

注:_remote_files是流式数据集样本的隐含字段,存储了该样本对应文件的远程访问地址。

3. 优化流式缓存策略确保文件完整缓存

通过指定缓存目录,触发音频数据的自动缓存后再加载:

# 加载数据集时指定缓存目录
streaming_dataset = load_dataset(
    "mozilla-foundation/common_voice_17_0",
    language,
    split="train",
    use_auth_token=True,
    trust_remote_code=True,
    streaming=True,
    cache_dir="./hf_dataset_cache"  # 自定义缓存目录
)

# 遍历数据集时访问audio字段触发缓存,找到目标样本后加载
target_idx = 0  # 替换为你需要的样本索引
for idx, sample in enumerate(dataset):
    # 访问audio字段会自动下载并缓存音频数据
    audio_data = sample["audio"]
    if idx == target_idx:
        waveform = torch.tensor(audio_data["array"])
        sample_rate = audio_data["sampling_rate"]
        break

内容的提问来源于stack exchange,提问作者Bobby Miller

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 04:18:19