如何统计数据集中各字符串的重复出现次数?
字符串重复次数统计解决方案
方案一:使用awk命令(Linux/macOS终端)
直接提取每行的目标字符串,用关联数组统计出现次数,最后遍历输出结果:
awk '{count[$2]++} END {for (str in count) print str, count[str]}' input.txt
如果无需文件,直接在终端输入数据处理:
cat << EOF | awk '{count[$2]++} END {for (str in count) print str, count[str]}' 1 AB 2 ZZ 3 BC 4 AB 5 ZZ 6 CC EOF
方案二:使用sort+uniq组合命令
通过多步命令协作完成统计,适合熟悉基础终端工具的场景:
cut -d' ' -f2 input.txt | sort | uniq -c | awk '{print $2, $1}'
各步骤说明:
cut -d' ' -f2:以空格为分隔符,提取每行第二个字段sort:对提取的字符串排序(uniq仅能统计连续重复的内容)uniq -c:统计每个字符串的出现次数,输出格式为「次数 字符串」awk '{print $2, $1}':调换输出顺序,转为「字符串 次数」
方案三:Python脚本实现
适合需要自定义逻辑或跨平台场景:
from collections import defaultdict # 从文件读取数据 count = defaultdict(int) with open('input.txt', 'r') as f: for line in f: # 分割每行并取最后一个元素(兼容任意数量的前缀空格/数字) target_str = line.strip().split()[-1] count[target_str] += 1 # 输出统计结果 for s, num in count.items(): print(f"{s} {num}")
若直接在脚本内嵌入数据:
from collections import defaultdict data = """1 AB 2 ZZ 3 BC 4 AB 5 ZZ 6 CC""" count = defaultdict(int) for line in data.splitlines(): target_str = line.strip().split()[-1] count[target_str] += 1 for s, num in count.items(): print(f"{s} {num}")
内容的提问来源于stack exchange,提问作者Reda
相关产品推荐
相关产品推荐

