如何统计Pandas中指定列字符串在另一列表列的出现次数?
Pandas列字符串在列表列中的出现次数统计
方法一:逐行遍历统计(符合需求)
先构造示例数据,再通过iterrows()遍历每一行,统计目标字符串在对应列表列中的出现次数,最后累加得到每个类别的总次数:
import pandas as pd # 示例数据(根据实际场景调整格式) data = { 'target_category': [ "Category A. Molecular Pathogenesis and Physiology", "Category B. Diagnosis and Assessment", "Category C. Treatment", "Category A. Molecular Pathogenesis and Physiology" ], 'top_predicted': [ ["Category A", "Category B"], ["Category B", "Category C", "Category B"], ["Category A", "Category C"], ["Category C", "Category D"] ] } df = pd.DataFrame(data) # 初始化计数字典 category_counts = {} # 遍历每一行统计 for _, row in df.iterrows(): target = row['target_category'] # 若target与top_predicted中的字符串格式一致,直接用target;否则提取匹配部分 match_str = target.split('.')[0].strip() current_count = row['top_predicted'].count(match_str) # 更新计数 if target in category_counts: category_counts[target] += current_count else: category_counts[target] = current_count # 将计数结果合并到原DataFrame df['occurrence_count'] = df['target_category'].map(category_counts) print(df)
方法二:用apply简化遍历
如果觉得循环代码冗余,可以用apply结合分组求和实现:
def count_match(row): match_str = row['target_category'].split('.')[0].strip() return row['top_predicted'].count(match_str) # 按目标类别分组,统计每组的总出现次数 total_counts = df.groupby('target_category').apply(lambda g: g.apply(count_match, axis=1).sum()) # 合并结果到原DataFrame df = df.merge(total_counts.rename('occurrence_count'), on='target_category', how='left')
关键注意事项
- 如果你的
target_category和top_predicted中的字符串完全一致,直接删除split('.')[0].strip()这一步,用target作为匹配字符串即可。 - 上述代码会统计每个目标类别在所有行的
top_predicted列表中的总出现次数,满足遍历所有行的需求。
内容的提问来源于stack exchange,提问作者dinho
相关产品推荐
相关产品推荐

