You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将字符串格式字典转为扁平字典并展平DataFrame列

问题描述

尝试展平DataFrame中的my_df['replies']列时,使用pd.json_normalize报错'str' object has no attribute 'values',现有解决方案无效。目标是生成包含ID列(对应多条评论会重复)、comment_id列及其他所有键对应列的DataFrame。

可复现的DataFrame示例:

from datetime import datetime
import pandas as pd
from pandas import DataFrame

my_dict = {'ID': {166: 166, 167: 167},
 'replies': {166: "[{'comment_id': '543806557410845', 'comment_url': 'https://facebook.com/543806557410845', 'commenter_id': '100044188131662', 'commenter_url': 'https://facebook.com/dtop?fref=nf&rc=p&__tn__=RR', 'commenter_name': 'Departamento de Transportación y Obras Públicas', 'commenter_meta': 'Author', 'comment_text': 'Mariangelies Morales Rodriguez si ya se ha hecho el trasaso en COlecturia el comprado debe ir al CESCO a registrarlo a su bombre.', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}, {'comment_id': '750421079486944', 'comment_url': 'https://facebook.com/750421079486944', 'commenter_id': '100000032881612', 'commenter_url': 'https://facebook.com/mariangelies.moralesrodriguez1?fref=nf&rc=p&__tn__=R', 'commenter_name': 'Mariangelies Morales Rodriguez', 'commenter_meta': None, 'comment_text': 'Departamento de Transportación y Obras Públicas gracias', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}]",
  167: "[{'comment_id': '806816950348709', 'comment_url': 'https://facebook.com/806816950348709', 'commenter_id': '100044188131662', 'commenter_url': 'https://facebook.com/dtop?fref=nf&rc=p&__tn__=RR', 'commenter_name': 'Departamento de Transportación y Obras Públicas', 'commenter_meta': 'Author', 'comment_text': 'Jorge Moran Suarez Usted debe solicitar al vendedor un certificado de multas con la fecha del traspaso, eso le indica si tiene multas en el sistema.', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}, {'comment_id': '1422909098136997', 'comment_url': 'https://facebook.com/1422909098136997', 'commenter_id': '100066536146341', 'commenter_url': 'https://facebook.com/profile.php?id=100066536146341&fref=nf&rc=p&__tn__=R', 'commenter_name': 'Jose Serrano', 'commenter_meta': None, 'comment_text': 'Departamento de Transportación y Obras Públicas cual es el proceso para dar de baja unos vehiculos q estan a mi nombre y q solo tengo los num.de tablilla', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}]"}}  
my_df = DataFrame(my_dict)
解决方案

报错核心原因是replies列的内容是字符串格式的列表,而非实际的Python列表对象,pd.json_normalize无法直接解析字符串。解决步骤如下:

  1. 将字符串转成Python列表:用ast.literal_eval解析字符串(示例中的单引号语法合法,ast可直接处理)。
  2. 拆分评论行并关联ID:用explode把每个ID对应的多条评论拆分成单独行,再用json_normalize展平评论字典。

完整代码:

import ast
import pandas as pd

# 1. 解析replies列的字符串为Python列表
my_df['replies'] = my_df['replies'].apply(ast.literal_eval)

# 2. 将每个ID对应的多条评论拆分为单独行
exploded_df = my_df.explode('replies', ignore_index=True)

# 3. 展平评论字典,同时保留原ID列
final_df = pd.concat([
    exploded_df[['ID']],
    pd.json_normalize(exploded_df['replies'])
], axis=1)

# 查看结果
print(final_df)
结果说明

最终生成的DataFrame会包含:

  • 重复的ID列(每个ID对应每条评论)
  • comment_id、comment_url、commenter_id等所有评论字段列

内容的提问来源于stack exchange,提问作者Steven González

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 23:35:20