You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计多CSV文件邮件列中的重复句子?

多CSV邮件重复句子统计方案

问题

有多份CSV文件,其中一列存储完整邮件内容(每行对应一封邮件),需遍历所有邮件统计重复出现的句子。若仅用pandas,要求每行对应一个句子,但当前数据中每封邮件包含多个句子。

输入输出示例

输入文本

This is a new first line.
This line is being added to file.
And another line here.

This is more text being added to file.
And another line here.
This is a new first line

预期输出

SentenceCount
This is a new first line2
And another line here.2
This line is being added to file.1
This is more text being added to file.1

当前问题

使用spaCy进行句子分割时,会将跨换行的句子合并,导致重复句子无法被正确统计。

当前代码及输出

句子分割代码

# 用spaCy获取句子数量
sentence_tokens = [[token.text for token in sent] for sent in doc.sents]
print(len(sentence_tokens))

# 提取句子列表
[sent.text for sent in doc.sents ]

输出:

5

['This is a new first line.\n',
 'This line is being appended to file\nAnd another line here.\n',
 'This is more text being appended to file.\n',
 'And another line here.\n',
 'This is a new first line\n']

计数代码

from collections import Counter
counts = Counter([sent.text for sent in doc.sents ])
print(counts)

输出:

Counter({'This is a new first line.\n': 1, 'This line is being appended to file\nAnd another line here.\n': 1, 'This is more text being appended to file.\n': 1, 'And another line here.\n': 1, 'This is a new first line\n': 1})

解决方案

核心思路

先按换行符拆分邮件内容,过滤空行后再逐行进行句子分割,避免spaCy跨行合并句子;同时标准化句子格式(清理首尾空白、统一标点),确保相同语义的句子被正确识别为重复。

完整实现代码

import pandas as pd
import spacy
from collections import Counter

# 加载spaCy英文模型
nlp = spacy.load("en_core_web_sm")

# 替换为你的CSV文件路径列表
csv_files = ["email_1.csv", "email_2.csv"]
all_cleaned_sentences = []

# 遍历所有CSV文件
for file_path in csv_files:
    df = pd.read_csv(file_path)
    # 替换为实际的邮件内容列名
    email_contents = df["email_content"].dropna().tolist()
    
    for content in email_contents:
        # 按换行拆分,过滤空行并清理每行首尾空白
        lines = [line.strip() for line in content.split("\n") if line.strip()]
        for line in lines:
            # 对单行文本进行句子分割
            doc = nlp(line)
            for sent in doc.sents:
                # 清理句子首尾空白,标准化格式
                cleaned_sent = sent.text.strip()
                all_cleaned_sentences.append(cleaned_sent)

# 统计句子出现次数
sent_count = Counter(all_cleaned_sentences)

# 转换为DataFrame并按次数降序排序
result_df = pd.DataFrame(sent_count.items(), columns=["Sentence", "Count"])
result_df = result_df.sort_values(by="Count", ascending=False).reset_index(drop=True)

# 输出Markdown格式结果
print(result_df.to_markdown(index=False))

代码说明

  1. CSV读取:批量读取所有目标CSV文件,提取邮件内容列并跳过空值。
  2. 换行拆分:将每封邮件按换行拆分为独立行,过滤空行避免无效处理。
  3. 句子分割与标准化:对每行文本单独做句子分割,清理句子首尾空白,保证相同句子格式统一。
  4. 统计与输出:用Counter统计句子频次,转换为DataFrame后排序,输出为Markdown表格。

内容的提问来源于stack exchange,提问作者H-Finch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 13:18:09