You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何TensorFlow_IO AudioIOTensor多实例读取结果重复?如何同时使用?

问题

我正在学习如何在TensorFlow中处理音频数据,已下载Kaggle Speech Accent Archive数据集。尝试对比两种TensorFlow音频处理方式:

  • 第一种:通过tf.io.read_file读取二进制文件,再调用tfio.audio.decode_mp3解码MP3文件
  • 第二种:将文件读取为AudioIOTensor,它会自动识别MP3格式并存储元数据,采用延迟加载机制,仅在明确调用时才加载全部数据

用于对比的代码如下:

import pathlib
import tensorflow as tf
import tensorflow_io as tfio
DATA_DIR = <Path to data>

data_path = pathlib.Path(DATA_DIR)
mp3Files = [x for x in data_path.iterdir() if '.mp3' in x.name]

def load_audios(file_list):
    dataset = []
    for curr_file in file_list:
        curr_file_path = curr_file.as_posix()
        audio_binary = tf.io.read_file(curr_file_path)
        audio1 = tfio.audio.decode_mp3(audio_binary)
        audio2 = tfio.audio.AudioIOTensor(curr_file_path)
        dataset.append([audio1,audio2, curr_file.name])
    return dataset

test = load_audios(mp3Files[0:3])

运行后成功读取三个文件:

>>> test[0][-1], test[1][-1], test[2][-1]
('afrikaans1.mp3', 'afrikaans2.mp3', 'afrikaans3.mp3')

但出现异常:调用test[0][1].to_tensor()、test[1][1].to_tensor()和test[2][1].to_tensor()得到的结果完全一致,且与test[2][0](最后一个文件的decode_mp3解码结果)匹配;而test[0][0]和test[1][0]分别对应前两个文件的正确解码结果。

请问该异常的原因是什么?如何才能同时使用多个不同的AudioIOTensor实例?

原因分析

问题出在TensorFlow的计算图追踪机制与AudioIOTensor的延迟加载特性的交互上:

  • 在Python循环中,curr_file_path是一个不断被重新赋值的Python字符串变量
  • TensorFlow在创建AudioIOTensor时,会自动追踪这个变量作为计算图的一部分,而非将每个循环迭代的路径值固化
  • 当后续调用to_tensor()触发实际加载时,所有AudioIOTensor实例都会使用循环结束时最后一次赋值的curr_file_path,导致全部加载最后一个文件的内容
解决方法

要为每个AudioIOTensor实例绑定独立的文件路径,需要将Python字符串路径转换为TensorFlow的常量张量,避免被计算图追踪复用:

修改循环内创建AudioIOTensor的代码,将路径转为tf.string常量:

def load_audios(file_list):
    dataset = []
    for curr_file in file_list:
        curr_file_path = curr_file.as_posix()
        audio_binary = tf.io.read_file(curr_file_path)
        audio1 = tfio.audio.decode_mp3(audio_binary)
        # 将路径转为TensorFlow常量张量
        tensor_path = tf.convert_to_tensor(curr_file_path, dtype=tf.string)
        audio2 = tfio.audio.AudioIOTensor(tensor_path)
        dataset.append([audio1,audio2, curr_file.name])
    return dataset

这样每个AudioIOTensor实例都会持有独立的路径张量,调用to_tensor()时会加载对应文件的内容,不会再出现复用最后一个文件的问题。

内容的提问来源于stack exchange,提问作者user1245262

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 13:10:35