You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计数据集中各字符串的重复出现次数?

字符串重复次数统计解决方案

方案一:使用awk命令(Linux/macOS终端)

直接提取每行的目标字符串,用关联数组统计出现次数,最后遍历输出结果:

awk '{count[$2]++} END {for (str in count) print str, count[str]}' input.txt

如果无需文件,直接在终端输入数据处理:

cat << EOF | awk '{count[$2]++} END {for (str in count) print str, count[str]}'
1      AB
2      ZZ
3      BC
4      AB
5      ZZ
6      CC
EOF

方案二:使用sort+uniq组合命令

通过多步命令协作完成统计,适合熟悉基础终端工具的场景:

cut -d' ' -f2 input.txt | sort | uniq -c | awk '{print $2, $1}'

各步骤说明:

  • cut -d' ' -f2:以空格为分隔符,提取每行第二个字段
  • sort:对提取的字符串排序(uniq仅能统计连续重复的内容)
  • uniq -c:统计每个字符串的出现次数,输出格式为「次数 字符串」
  • awk '{print $2, $1}':调换输出顺序,转为「字符串 次数」

方案三:Python脚本实现

适合需要自定义逻辑或跨平台场景:

from collections import defaultdict

# 从文件读取数据
count = defaultdict(int)
with open('input.txt', 'r') as f:
    for line in f:
        # 分割每行并取最后一个元素(兼容任意数量的前缀空格/数字)
        target_str = line.strip().split()[-1]
        count[target_str] += 1

# 输出统计结果
for s, num in count.items():
    print(f"{s} {num}")

若直接在脚本内嵌入数据:

from collections import defaultdict

data = """1      AB
2      ZZ
3      BC
4      AB
5      ZZ
6      CC"""

count = defaultdict(int)
for line in data.splitlines():
    target_str = line.strip().split()[-1]
    count[target_str] += 1

for s, num in count.items():
    print(f"{s} {num}")

内容的提问来源于stack exchange,提问作者Reda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 06:35:26