统计Pandas列中含字典列表的单元格内字典元素数量
Got it, let's break down how to count the number of dictionary elements in each cell of your df['likes'] column. The column has a mix of NaNs, and cells containing dictionaries with a data key that holds a list of dictionaries—we need to grab the length of that data list for each valid cell.
方法一:自定义函数处理(清晰直观)
First, we can write a small helper function to handle each cell, then apply it to the column:
import pandas as pd def count_like_entries(cell): # 先处理空值 if pd.isna(cell): return 0 # 如果你想保留空值标记,可以换成 pd.NA # 检查单元格是否是包含'data'键的有效字典 if isinstance(cell, dict) and 'data' in cell: return len(cell['data']) # 针对不符合预期格式的情况(非字典或缺少'data'键)返回默认值 return 0 # 生成统计列 df['like_entry_count'] = df['likes'].apply(count_like_entries)
方法二:简洁的Lambda表达式
If you prefer a more compact approach, you can combine str.get() with a lambda to skip writing a separate function:
# 先提取'data'列表,再统计长度(同时处理空值) df['like_entry_count'] = df['likes'].str.get('data').apply(lambda x: len(x) if not pd.isna(x) else 0)
额外处理:如果字典是字符串格式
In case some cells have dictionary-like strings (instead of actual Python dictionaries), you'll need to convert them first using ast.literal_eval:
import ast # 将字符串格式的字典转换为真实的Python字典对象 df['likes'] = df['likes'].apply(lambda x: ast.literal_eval(x) if isinstance(x, str) else x) # 然后运行上面的统计代码 df['like_entry_count'] = df['likes'].str.get('data').apply(lambda x: len(x) if not pd.isna(x) else 0)
两种方法都会生成一个新列,里面是每个likes单元格内的字典元素数量——空值单元格会显示0(你也可以根据需求调整为显示pd.NA来保留空值标记)。
内容的提问来源于stack exchange,提问作者Edward

