You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas append追加列表数据至DataFrame时出现NaN值问题

解决Pandas DataFrame追加数据全为NaN的问题

我来帮你排查下问题所在,你的代码里有几个关键问题导致了所有列都显示NaN:

问题分析

  1. 循环变量冲突:外层循环用了i作为计数器,内层循环又重复使用了i,这会直接覆盖外层的i值,导致循环逻辑混乱,完全达不到分批处理的效果。
  2. 错误的元素引用:你在内层循环里一直取dumps[0].values(),相当于每次都在添加批次里第一个音频特征的数据,而不是当前遍历的元素。
  3. 列名不匹配:初始的features_df是空DataFrame,当你直接append一个列表时,Pandas会把列表元素对应到索引为0、1、2...的列上,而你期望的danceability、energy等列根本不存在,所以所有目标列都会显示NaN。

修正后的解决方案

推荐先收集所有音频特征字典,再一次性转换为DataFrame(这种方式效率更高,避免多次append的性能损耗):

import pandas as pd
import spotipy as sp

# 初始化列表存储所有有效音频特征字典
all_features = []

# 分批处理ids
for batch_idx in range(int(len(ids)/50)):
    # 截取当前批次的id
    batch_ids = ids[batch_idx*50 : (batch_idx+1)*50]
    # 获取当前批次的音频特征
    batch_features = sp.audio_features(batch_ids)
    # 过滤掉请求失败返回的None值(避免后续转换出错)
    valid_features = [feat for feat in batch_features if feat is not None]
    # 将有效特征加入总列表
    all_features.extend(valid_features)

# 直接将字典列表转换为DataFrame,自动匹配列名
features_df = pd.DataFrame(all_features)

如果你坚持要逐行追加数据(不推荐,大数据量下性能差),可以修改为以下代码,确保每次追加的是完整的字典而非列表:

import pandas as pd
import spotipy as sp

features_df = pd.DataFrame()

for batch_idx in range(int(len(ids)/50)):
    batch_ids = ids[batch_idx*50 : (batch_idx+1)*50]
    batch_features = sp.audio_features(batch_ids)
    # 内层循环用j作为计数器,避免和外层batch_idx冲突
    for j in range(len(batch_features)):
        current_feat = batch_features[j]
        if current_feat is not None:
            # 追加字典,Pandas会自动匹配已有的列名(首次追加时自动创建列)
            features_df = features_df.append(current_feat, ignore_index=True)

关键说明

  • 音频特征接口可能会返回None(比如id无效或请求超时),所以一定要过滤掉这些空值,避免转换DataFrame时出错。
  • 直接使用字典列表创建DataFrame是Pandas中最高效的方式,比逐行append性能提升很多,尤其是处理大量数据时。

内容的提问来源于stack exchange,提问作者the.lotuseater

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:52:11