如何为Pandas DataFrame新增列统计列表元素出现次数并生成字典?
解决Pandas统计字符串元素出现次数并生成字典列的问题
我来帮你搞定这个问题!你之前用str.count(k)的方式行不通,是因为这个方法只能针对单个指定字符串统计次数,没法动态对每行的所有元素批量统计。我们可以结合apply和collections.Counter来实现需求,具体步骤如下:
步骤1:导入必要库并定义原始DataFrame
import pandas as pd from collections import Counter # 你的原始DataFrame df = pd.DataFrame({'StringMatch':['[mother]','[priest, mother,mother father]','[father, mother]']})
步骤2:编写元素统计函数
这个函数会处理每行的StringMatch字符串,完成清理、分割、统计三个核心操作:
def calculate_element_counts(s): # 去掉字符串首尾的方括号 cleaned_str = s.strip('[]') # 按逗号分割元素,同时去掉每个元素前后的空格(处理逗号后无空格的情况) elements = [elem.strip() for elem in cleaned_str.split(',')] # 统计每个元素的出现次数,转成字典 count_result = dict(Counter(elements)) # 转换成你需要的字符串格式(如果要存储字典对象,直接返回count_result即可) return '{' + ', '.join([f"{key}:{value}" for key, value in count_result.items()]) + '}'
步骤3:应用函数生成新列
把函数应用到StringMatch列上,生成StringMatchCount列:
df['StringMatchCount'] = df['StringMatch'].apply(calculate_element_counts)
最终结果
执行后你的DataFrame会变成:
| StringMatch | StringMatchCount |
|---|---|
| [mother] | {mother:1} |
| [priest, mother,mother father] | {priest:1, mother:1, mother father:1} |
| [father, mother] | {father:1, mother:1} |
哦对了,看你给出的目标示例里第二个结果是{priest:1, mother:2, father:1},猜测你原始的第二个字符串可能是笔误(应该是[priest, mother, mother, father]),如果是这样的话,上面的代码会自动统计出mother:2的结果,完全符合你的预期。
如果后续需要对统计结果做数值运算,建议直接存储字典对象(去掉函数里的字符串格式化步骤,返回count_result即可),这样可以直接通过键获取次数值,比操作字符串更灵活。
内容的提问来源于stack exchange,提问作者wwnde
相关产品推荐
相关产品推荐

