You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas统计DataFrame每行指定列的逗号分隔唯一字符串数量

pandas实现按逗号分割统计每行唯一值数量

实现思路

对type列每行的字符串先做空值判断,再按逗号分割、去除元素前后空格、过滤无效空字符串后去重计数,空值直接返回0。

完整实现代码

首先导入依赖库:

import pandas as pd
import numpy as np

构造示例DataFrame(如果已有数据可跳过这一步):

df = pd.DataFrame({
    'no': [1, 2, 3, 4],
    'type': ['123, 234, 345', '123', '4567,235', np.nan]
})

写法1:自定义函数(可读性更高)

def count_unique(s):
    if pd.isna(s):
        return 0
    # 分割、去空格、过滤空值、去重计数
    items = [item.strip() for item in s.split(',')]
    valid_unique = set(filter(lambda x: x, items))
    return len(valid_unique)

df['count'] = df['type'].apply(count_unique)

写法2:一行式实现

df['count'] = df['type'].fillna('').str.split(',').apply(lambda x: len(set([i.strip() for i in x if i.strip()])))

输出结果验证

执行后输出df即可得到预期结果:

no           type  count
0   1  123, 234, 345      3
1   2            123      1
2   3       4567,235      2
3   4            NaN      0

内容的提问来源于stack exchange,提问作者nomnom3214

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 04:45:03