You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计Pandas中指定列字符串在另一列表列的出现次数?

Pandas列字符串在列表列中的出现次数统计

方法一:逐行遍历统计(符合需求)

先构造示例数据,再通过iterrows()遍历每一行,统计目标字符串在对应列表列中的出现次数,最后累加得到每个类别的总次数:

import pandas as pd

# 示例数据(根据实际场景调整格式)
data = {
    'target_category': [
        "Category A. Molecular Pathogenesis and Physiology",
        "Category B. Diagnosis and Assessment",
        "Category C. Treatment",
        "Category A. Molecular Pathogenesis and Physiology"
    ],
    'top_predicted': [
        ["Category A", "Category B"],
        ["Category B", "Category C", "Category B"],
        ["Category A", "Category C"],
        ["Category C", "Category D"]
    ]
}
df = pd.DataFrame(data)

# 初始化计数字典
category_counts = {}

# 遍历每一行统计
for _, row in df.iterrows():
    target = row['target_category']
    # 若target与top_predicted中的字符串格式一致,直接用target;否则提取匹配部分
    match_str = target.split('.')[0].strip()
    current_count = row['top_predicted'].count(match_str)
    
    # 更新计数
    if target in category_counts:
        category_counts[target] += current_count
    else:
        category_counts[target] = current_count

# 将计数结果合并到原DataFrame
df['occurrence_count'] = df['target_category'].map(category_counts)
print(df)

方法二:用apply简化遍历

如果觉得循环代码冗余,可以用apply结合分组求和实现:

def count_match(row):
    match_str = row['target_category'].split('.')[0].strip()
    return row['top_predicted'].count(match_str)

# 按目标类别分组,统计每组的总出现次数
total_counts = df.groupby('target_category').apply(lambda g: g.apply(count_match, axis=1).sum())

# 合并结果到原DataFrame
df = df.merge(total_counts.rename('occurrence_count'), on='target_category', how='left')

关键注意事项

  • 如果你的target_category和top_predicted中的字符串完全一致,直接删除split('.')[0].strip()这一步,用target作为匹配字符串即可。
  • 上述代码会统计每个目标类别在所有行的top_predicted列表中的总出现次数,满足遍历所有行的需求。

内容的提问来源于stack exchange,提问作者dinho

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 09:15:22