流式加载HuggingFace数据集时音频文件无法下载加载求助
流式加载Common Voice数据集时音频加载失败的解决方案
问题描述
流式加载HuggingFace数据集mozilla-foundation/common_voice_17_0的代码如下:
language = "en" buffer_size = 100 streaming_dataset = load_dataset("mozilla-foundation/common_voice_17_0", language, split="train", use_auth_token=True, trust_remote_code=True, streaming=True) dataset = streaming_dataset.shuffle( seed=1137, buffer_size=buffer_size ) common_voice = buffer_dataset(dataset, buffer_size) sample = random.choice(common_voice) audio_file = sample["path"]
所有步骤执行正常,但调用torchaudio.load(audio_file)加载音频时出现报错:
Traceback (most recent call last): ... "... .venv/lib/python3.10/site-packages/soundfile.py", line 1216, in _open raise LibsndfileError(err, prefix="Error opening {0!r}: ".format(self.name)) soundfile.LibsndfileError: Error opening 'en_train_14/common_voice_en_18837621.mp3': System error.
非流式加载小数据集分片可正常运行,但'en'这类大数据集无法采用该方案,推测是streaming=True时音频文件未完成下载导致的问题。
解决方案
1. 直接使用数据集自带的音频字段加载
流式模式下,样本的audio字段已包含音频的原始数据和采样率,无需依赖本地文件路径加载:
# 从sample的audio字段直接提取波形和采样率 waveform = torch.tensor(sample["audio"]["array"]) sample_rate = sample["audio"]["sampling_rate"]
这种方式会自动流式获取音频数据,无需等待完整文件下载,完全适配流式场景。
2. 手动下载指定音频文件(适用于必须用本地文件的场景)
如果需要将音频文件保存到本地后再加载,可以利用样本隐含的远程文件信息,通过数据集构建器单独下载:
from datasets import load_dataset_builder # 获取数据集构建器,用于处理文件下载 ds_builder = load_dataset_builder("mozilla-foundation/common_voice_17_0", language) # 获取当前样本对应的远程音频文件URL remote_file_url = sample["_remote_files"][0]["url"] # 指定本地保存路径 local_file_path = "./target_audio.mp3" # 下载文件到本地 ds_builder.download(remote_file_url, local_file_path) # 现在可以正常加载本地文件 waveform, sample_rate = torchaudio.load(local_file_path)
注:_remote_files是流式数据集样本的隐含字段,存储了该样本对应文件的远程访问地址。
3. 优化流式缓存策略确保文件完整缓存
通过指定缓存目录,触发音频数据的自动缓存后再加载:
# 加载数据集时指定缓存目录 streaming_dataset = load_dataset( "mozilla-foundation/common_voice_17_0", language, split="train", use_auth_token=True, trust_remote_code=True, streaming=True, cache_dir="./hf_dataset_cache" # 自定义缓存目录 ) # 遍历数据集时访问audio字段触发缓存,找到目标样本后加载 target_idx = 0 # 替换为你需要的样本索引 for idx, sample in enumerate(dataset): # 访问audio字段会自动下载并缓存音频数据 audio_data = sample["audio"] if idx == target_idx: waveform = torch.tensor(audio_data["array"]) sample_rate = audio_data["sampling_rate"] break
内容的提问来源于stack exchange,提问作者Bobby Miller
相关产品推荐
相关产品推荐

