You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用literal_eval提取流派名称遇“malformed node or string”错误求助

修复方案与更优提取方法

错误原因

ast.literal_eval报错"malformed node or string",大概率是genres列存在格式不规范的字符串:

  • 用单引号代替了JSON要求的双引号
  • 存在多余的尾逗号
  • 除了填充的'[]'外还有其他非法格式的内容

修复后的代码

改用json.loads替代literal_eval,配合异常处理兼容格式问题:

import json
import pandas as pd

def extract_genres(genres_str):
    try:
        # 处理可能存在的单引号转双引号问题
        fixed_str = genres_str.replace("'", '"') if isinstance(genres_str, str) else genres_str
        genres_list = json.loads(fixed_str)
        return [item['name'] for item in genres_list] if isinstance(genres_list, list) else None
    except (json.JSONDecodeError, TypeError, KeyError):
        return None

# 应用到目标列
df['genres'] = df['genres'].fillna('[]').apply(extract_genres)

如果你的数据是标准JSON格式(全双引号),可简化为:

df['genres'] = df['genres'].fillna('[]').apply(json.loads).apply(lambda x: [i['name'] for i in x] if isinstance(x, list) else None)

更优提取方案

方案1:用pd.json_normalize展开数据

适合需要拆分流派到单独行或批量提取的场景:

# 先将字符串转为列表格式
df['genres_list'] = df['genres'].fillna('[]').apply(json.loads)
# 展开列表并提取流派名称
genres_df = pd.json_normalize(df['genres_list']).melt(value_name='genre').dropna()
# 合并回原表(按需选择)
df['genres'] = df['genres_list'].apply(lambda x: [g['name'] for g in x])

方案2:正则提取(高效适配格式稳定的数据集)

如果genres列字符串格式统一,直接用正则匹配name字段值,速度更快:

df['genres'] = df['genres'].str.findall(r'"name":"([^"]+)"').apply(lambda x: x if x else None)

内容的提问来源于stack exchange,提问作者william_Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 03:42:18