You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算DataFrame字符串列中分类标签的Value Counts?

问题描述

现有一个包含字符串列col的DataFrame,数据示例如下:

df col
0 fruit["apple"], colour["green", "yellow"] 
1 colour["yellow"] 
2 colour["brown"] 

需要计算该列中各类标签(如fruit、colour)的出现次数,预期输出结果如下:

fruit  1
colour 3
解决方案

使用pandas结合正则表达式即可实现需求,具体代码如下:

import pandas as pd
import re

# 构造示例数据
data = {'col': [
    'fruit["apple"], colour["green", "yellow"]',
    'colour["yellow"]',
    'colour["brown"]'
]}
df = pd.DataFrame(data)

# 正则匹配标签(匹配["前的单词)
tag_pattern = re.compile(r'(\w+)\["')

# 提取每行的所有标签并展开为单行一个标签
all_tags = df['col'].str.findall(tag_pattern).explode()

# 统计各标签出现次数
tag_count_result = all_tags.value_counts()

print(tag_count_result)

运行代码后输出:

colour    3
fruit     1
dtype: int64

内容的提问来源于stack exchange,提问作者arv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 12:35:28