You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用LSTM进行音频情感分析时遭遇NumPy转Tensor错误求助

LSTM音频情感分析:解决NumPy转Tensor错误问题

我正在尝试用LSTM做基于音频文件的情感分析,输入包含Sentiment(取值为Negative、Positive、Neutral)和soundFile(音频文件)。数据共28行,训练集16行,每行包含101个值。试过网上各种方案,还是卡在这个错误:

ValueError: Failed to convert a NumPy array to a Tensor (Unsupported object type numpy.ndarray)

相关代码如下:

#Extracts the sound features
def extract_sound_feature(audio_file, max_length=100):
    audio, sr=librosa.load(audio_file, sr=None)
    mfccs=librosa.feature.mfcc(y=audio, sr=sr, n_mfcc=20)
    mfccs_normalized=(mfccs - mfccs.mean()) / mfccs.std()
    mfccs_padded = pad_sequences([mfccs_normalized], maxlen=max_length, padding='post', truncating='post')[0]
    return mfccs_padded
#Importing sound files
def import_sound_files(table_df, sound_df, folder_name):
    for i in range(len(table_df)):
       file_name = f"{folder_name}/{table_df['fileName'][i]}"
       features = extract_sound_feature(file_name, 20)
       temp_df=pd.DataFrame({
           'fileName': [table_df['fileName'][i]],
           'soundFile': [features]
           }) 
       sound_df=pd.concat([sound_df, temp_df], ignore_index=True)
    return sound_df
#Training LSTM model
def lstm_model_training(table_df):
    le_x = LabelEncoder()
    le_y = LabelEncoder()
    table_df['Sentiment'] = le_x.fit_transform(table_df['Sentiment'])
    table_df['Emotion'] = le_y.fit_transform(table_df['Emotion'])

    #Adding "Sentiment" and "soundFile" column together
    table_df['soundFile'] = np.array(table_df['soundFile'])
    for i in range(len(table_df['soundFile'])):
        table_df['soundFile'][i] = np.append(table_df['soundFile'][i], table_df['Sentiment'][i])
    
    x=table_df['soundFile']
    y=np.array(table_df['Emotion'])

    
    x_train, x_test, y_train, y_test= train_test_split(x, y, test_size=0.25, random_state=42)
    x_train, x_validation, y_train, y_validation= train_test_split(x_train, y_train, test_size=0.2)
    
    input_shape = (101, 1)  # Correct input shape
    model= keras.Sequential()
    #Adding layers
    model.add(keras.layers.LSTM(64, input_shape=input_shape, return_sequences=True))

    #Dense layers
    model.add(keras.layers.Dense(64, activation= 'relu'))
    model.add(keras.layers.Dropout(0.3))
    
    #Output layer
    model.add(keras.layers.Dense(10, activation='softmax'))

    #Compiling 
    optimiser = keras.optimizers.Adam(learning_rate=0.0001)
    model.compile(optimizer=optimiser,
                  loss='sparse_categorical_crossentropy',
                  metrics=['accuracy'])

    model.summary()

    #Training
    model.fit(x_train, y_train, validation_data=(x_validation, y_validation), batch_size=32, epochs=30)
    loss, accuracy=model.evaluate(x_test, y_test, verbose=2)
    print('\nTest accuracy: ', accuracy)

    return model

问题根源

  1. 特征维度错误:MFCC特征是(n_mfcc, 时间步长)的二维数组,原代码直接对其padding后返回,后续拼接Sentiment时变成嵌套数组,TensorFlow无法解析。
  2. DataFrame存储问题:DataFrame的soundFile列存储单个numpy数组,转成np.array后变成object类型的数组集合,不符合TensorFlow的输入要求。
  3. 输入形状不匹配:LSTM需要(样本数, 时间步长, 特征数)的三维张量,原输入是一维的object数组,维度不兼容。

修正方案

1. 修正特征提取函数

将MFCC转置为(时间步长, 特征数)的结构,方便后续拼接全局特征:

def extract_sound_feature(audio_file, max_length=100):
    audio, sr = librosa.load(audio_file, sr=None)
    mfccs = librosa.feature.mfcc(y=audio, sr=sr, n_mfcc=20)
    mfccs_normalized = (mfccs - mfccs.mean()) / mfccs.std()
    # 转置为(时间步长, 特征数)的结构
    mfccs_transposed = mfccs_normalized.T
    # 填充到指定时间步长
    mfccs_padded = pad_sequences([mfccs_transposed], maxlen=max_length, padding='post', truncating='post')[0]
    return mfccs_padded

2. 修正数据拼接与张量转换

直接将特征转换为三维numpy数组,避免DataFrame嵌套存储问题,同时将Sentiment作为额外特征添加到每个时间步:

def lstm_model_training(table_df):
    le_x = LabelEncoder()
    le_y = LabelEncoder()
    table_df['Sentiment'] = le_x.fit_transform(table_df['Sentiment'])
    table_df['Emotion'] = le_y.fit_transform(table_df['Emotion'])

    # 将soundFile列转换为三维数组:(样本数, 时间步长, 20)
    x_features = np.array([arr for arr in table_df['soundFile'].values])
    # 将Sentiment扩展为(样本数, 时间步长, 1),匹配MFCC的时间步长
    sentiment_features = np.expand_dims(table_df['Sentiment'].values, axis=-1)
    sentiment_features = np.repeat(sentiment_features, x_features.shape[1], axis=1)
    # 拼接特征得到(样本数, 时间步长, 21)的三维张量
    x = np.concatenate([x_features, sentiment_features], axis=-1).astype('float32')
    y = table_df['Emotion'].values

    x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.25, random_state=42)
    x_train, x_validation, y_train, y_validation = train_test_split(x_train, y_train, test_size=0.2)
    
    # 修正输入形状为(时间步长, 总特征数)
    input_shape = (x_train.shape[1], x_train.shape[2])  
    model = keras.Sequential()
    model.add(keras.layers.LSTM(64, input_shape=input_shape, return_sequences=True))
    model.add(keras.layers.Dense(64, activation='relu'))
    model.add(keras.layers.Dropout(0.3))
    # 输出层维度需与Emotion类别数一致,若实际类别数不是10需调整
    model.add(keras.layers.Dense(10, activation='softmax'))

    optimiser = keras.optimizers.Adam(learning_rate=0.0001)
    model.compile(optimizer=optimiser,
                  loss='sparse_categorical_crossentropy',
                  metrics=['accuracy'])

    model.summary()

    # 此时x_train为标准三维张量,可被TensorFlow正确处理
    model.fit(x_train, y_train, validation_data=(x_validation, y_validation), batch_size=32, epochs=30)
    loss, accuracy = model.evaluate(x_test, y_test, verbose=2)
    print('\nTest accuracy: ', accuracy)

    return model

3. 修正导入函数的参数匹配

确保特征提取时的时间步长与训练时一致:

def import_sound_files(table_df, sound_df, folder_name):
    for i in range(len(table_df)):
       file_name = f"{folder_name}/{table_df['fileName'][i]}"
       # 传入与训练一致的时间步长100
       features = extract_sound_feature(file_name, max_length=100)
       temp_df = pd.DataFrame({
           'fileName': [table_df['fileName'][i]],
           'soundFile': [features]
       }) 
       sound_df = pd.concat([sound_df, temp_df], ignore_index=True)
    return sound_df

内容的提问来源于stack exchange,提问作者M. Burak Toker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 13:31:08