如何从含字典数组的df['comments']列生成text&score完整表格?
解决Pandas展开嵌套字典数组生成表格的问题
问题背景
我的DataFrame的comments列包含88107行数据,每行都是由字典组成的数组,示例结构如下:
[{"text": "It will be curious to see where this heads in the long run. CBS is on a tear but will it fit their image, will they try and establish control, overall agenda. I've enjoyed last.fm for many years supporting through paypal donations each time I expire...it'll be interesting.","score": 0}, {"text": "Does this mean that there's now a big-name company who will fight for the repeal of the recent streaming-music royalty hike?","score": 1}, {"text": 'Also on BBC News: http://news.bbc.co.uk/1/low/technology/6701863.stm .Nice to see a London-based co. hit the headlines.','score': 2}, {"text": "I don't understand what they do that is worth $70M a year. ","score": 3}, {"text": 'sold out too cheaply. given their leadership position, they should have ask for at least $500m','score': 4}]
我需要生成包含text和score列的完整表格,但之前写的循环代码只能处理单个数组,无法覆盖全部88107行:
l=pd.DataFrame() for i in range(len(df['comments'][0])+1): n1=pd.DataFrame(data=df['comments'][i]) l=pd.concat([l, n1], axis=0)
可行解决方案
方法1:explode + json_normalize(高效简洁,推荐)
用Pandas内置的explode把每行的数组拆成单独行,再用json_normalize将字典展开为列:
import pandas as pd # 拆分comments列的数组为单行 exploded_df = df.explode('comments', ignore_index=True) # 展开字典为text和score列 result = pd.json_normalize(exploded_df['comments'])
这种方法无需手动循环,处理大规模数据时性能远优于循环拼接。
方法2:列表推导式扁平化字典数组
先把所有行的字典收集到一个列表,再一次性转为DataFrame:
# 遍历所有行,收集所有字典到一个列表 all_comments = [item for sublist in df['comments'] for item in sublist] # 转换为目标DataFrame result = pd.DataFrame(all_comments)
这种方式也能避免循环拼接的低效问题,适合处理88107行的大规模数据。
原代码问题说明
你之前的循环只遍历了第一行数组的元素个数对应的范围,相当于只处理了前N行数据(N是第一行数组的长度),并没有遍历全部88107行的comments列内容,所以无法生成完整表格。
内容的提问来源于stack exchange,提问作者doriam
相关产品推荐
相关产品推荐

