You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从JSON提取数据后Python行重复计数功能失效问题排查

问题排查与解决方法

问题根源

你的word_count函数使用doc.split()方法,该方法默认按空格拆分文本,会把整行内容拆成单个单词,只能统计单词出现次数,而非整行的重复次数。同时原代码每次遇到目标行就单独调用统计函数,没有跨行收集所有目标行做聚合统计,导致输出的是每行的单词计数而非整行的重复次数。

解决步骤

1. 重构统计函数,按整行计数

修改函数逻辑,直接对完整行内容做统计,同时清理每行的空白字符、跳过空行:

def line_count(lines):
    counts = dict()
    for line in lines:
        # 清理行首尾的空白(换行符、多余空格等)
        cleaned_line = line.strip()
        # 跳过空行
        if not cleaned_line:
            continue
        # 统计整行出现次数
        if cleaned_line in counts:
            counts[cleaned_line] += 1
        else:
            counts[cleaned_line] = 1
    return counts

2. 调整主逻辑,先收集所有目标行再统一统计

原代码没有收集所有需要统计的行,而是逐行单独处理,修改后先把所有目标行存入列表,再一次性统计:

import operator

# 初始化列表,收集所有需要统计的行
source_lines = []

issues_test = s_json[['issues'][0]]
for item in issues_test:
    description = item['fields']["description"]
    for line in description.split("\n"):
        if "Source of detection" in line:
            # 提取并清理内容
            cleaned_source = line.strip("*Source of detection:*").strip()
            # 替换不需要的特殊字符(修正原代码的replace语法错误)
            cleaned_source = cleaned_source.replace("“", "_").replace("”", "_")
            source_lines.append(cleaned_source)

# 统一统计所有行的重复次数
counts = line_count(source_lines)
# 按出现次数降序排序
sorted_d = sorted(counts.items(), key=operator.itemgetter(1), reverse=True)
# 输出最终结果
for key, value in sorted_d:
    print(f"{key} :  {value}")

3. 关于预期结果的补充说明

你给出的预期结果中,将TIU reports and validations简化为TIU reports、Reports from CARE简化为Reports,如果需要这种截断式统计,可在清理行内容时添加截断逻辑,比如只保留前2个单词:

# 在cleaned_line = line.strip()之后添加
cleaned_line = " ".join(cleaned_line.split()[:2]).strip()

如果只是笔误,上面的基础代码会输出完整行的统计结果,与你的需求逻辑一致。


内容的提问来源于stack exchange,提问作者Felipe Useda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 03:48:23