如何将字符串格式字典转为扁平字典并展平DataFrame列
问题描述
尝试展平DataFrame中的my_df['replies']列时,使用pd.json_normalize报错'str' object has no attribute 'values',现有解决方案无效。目标是生成包含ID列(对应多条评论会重复)、comment_id列及其他所有键对应列的DataFrame。
可复现的DataFrame示例:
from datetime import datetime import pandas as pd from pandas import DataFrame my_dict = {'ID': {166: 166, 167: 167}, 'replies': {166: "[{'comment_id': '543806557410845', 'comment_url': 'https://facebook.com/543806557410845', 'commenter_id': '100044188131662', 'commenter_url': 'https://facebook.com/dtop?fref=nf&rc=p&__tn__=RR', 'commenter_name': 'Departamento de Transportación y Obras Públicas', 'commenter_meta': 'Author', 'comment_text': 'Mariangelies Morales Rodriguez si ya se ha hecho el trasaso en COlecturia el comprado debe ir al CESCO a registrarlo a su bombre.', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}, {'comment_id': '750421079486944', 'comment_url': 'https://facebook.com/750421079486944', 'commenter_id': '100000032881612', 'commenter_url': 'https://facebook.com/mariangelies.moralesrodriguez1?fref=nf&rc=p&__tn__=R', 'commenter_name': 'Mariangelies Morales Rodriguez', 'commenter_meta': None, 'comment_text': 'Departamento de Transportación y Obras Públicas gracias', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}]", 167: "[{'comment_id': '806816950348709', 'comment_url': 'https://facebook.com/806816950348709', 'commenter_id': '100044188131662', 'commenter_url': 'https://facebook.com/dtop?fref=nf&rc=p&__tn__=RR', 'commenter_name': 'Departamento de Transportación y Obras Públicas', 'commenter_meta': 'Author', 'comment_text': 'Jorge Moran Suarez Usted debe solicitar al vendedor un certificado de multas con la fecha del traspaso, eso le indica si tiene multas en el sistema.', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}, {'comment_id': '1422909098136997', 'comment_url': 'https://facebook.com/1422909098136997', 'commenter_id': '100066536146341', 'commenter_url': 'https://facebook.com/profile.php?id=100066536146341&fref=nf&rc=p&__tn__=R', 'commenter_name': 'Jose Serrano', 'commenter_meta': None, 'comment_text': 'Departamento de Transportación y Obras Públicas cual es el proceso para dar de baja unos vehiculos q estan a mi nombre y q solo tengo los num.de tablilla', 'comment_time': datetime.datetime(2022, 7, 8, 0, 0), 'comment_image': None, 'comment_reactors': [], 'comment_reactions': None, 'comment_reaction_count': None}]"}} my_df = DataFrame(my_dict)
解决方案
报错核心原因是replies列的内容是字符串格式的列表,而非实际的Python列表对象,pd.json_normalize无法直接解析字符串。解决步骤如下:
- 将字符串转成Python列表:用
ast.literal_eval解析字符串(示例中的单引号语法合法,ast可直接处理)。 - 拆分评论行并关联ID:用
explode把每个ID对应的多条评论拆分成单独行,再用json_normalize展平评论字典。
完整代码:
import ast import pandas as pd # 1. 解析replies列的字符串为Python列表 my_df['replies'] = my_df['replies'].apply(ast.literal_eval) # 2. 将每个ID对应的多条评论拆分为单独行 exploded_df = my_df.explode('replies', ignore_index=True) # 3. 展平评论字典,同时保留原ID列 final_df = pd.concat([ exploded_df[['ID']], pd.json_normalize(exploded_df['replies']) ], axis=1) # 查看结果 print(final_df)
结果说明
最终生成的DataFrame会包含:
- 重复的
ID列(每个ID对应每条评论) comment_id、comment_url、commenter_id等所有评论字段列
内容的提问来源于stack exchange,提问作者Steven González
相关产品推荐
相关产品推荐

