如何统计多CSV文件邮件列中的重复句子?
多CSV邮件重复句子统计方案
问题
有多份CSV文件,其中一列存储完整邮件内容(每行对应一封邮件),需遍历所有邮件统计重复出现的句子。若仅用pandas,要求每行对应一个句子,但当前数据中每封邮件包含多个句子。
输入输出示例
输入文本
This is a new first line. This line is being added to file. And another line here. This is more text being added to file. And another line here. This is a new first line
预期输出
| Sentence | Count |
|---|---|
| This is a new first line | 2 |
| And another line here. | 2 |
| This line is being added to file. | 1 |
| This is more text being added to file. | 1 |
当前问题
使用spaCy进行句子分割时,会将跨换行的句子合并,导致重复句子无法被正确统计。
当前代码及输出
句子分割代码
# 用spaCy获取句子数量 sentence_tokens = [[token.text for token in sent] for sent in doc.sents] print(len(sentence_tokens)) # 提取句子列表 [sent.text for sent in doc.sents ]
输出:
5 ['This is a new first line.\n', 'This line is being appended to file\nAnd another line here.\n', 'This is more text being appended to file.\n', 'And another line here.\n', 'This is a new first line\n']
计数代码
from collections import Counter counts = Counter([sent.text for sent in doc.sents ]) print(counts)
输出:
Counter({'This is a new first line.\n': 1, 'This line is being appended to file\nAnd another line here.\n': 1, 'This is more text being appended to file.\n': 1, 'And another line here.\n': 1, 'This is a new first line\n': 1})
解决方案
核心思路
先按换行符拆分邮件内容,过滤空行后再逐行进行句子分割,避免spaCy跨行合并句子;同时标准化句子格式(清理首尾空白、统一标点),确保相同语义的句子被正确识别为重复。
完整实现代码
import pandas as pd import spacy from collections import Counter # 加载spaCy英文模型 nlp = spacy.load("en_core_web_sm") # 替换为你的CSV文件路径列表 csv_files = ["email_1.csv", "email_2.csv"] all_cleaned_sentences = [] # 遍历所有CSV文件 for file_path in csv_files: df = pd.read_csv(file_path) # 替换为实际的邮件内容列名 email_contents = df["email_content"].dropna().tolist() for content in email_contents: # 按换行拆分,过滤空行并清理每行首尾空白 lines = [line.strip() for line in content.split("\n") if line.strip()] for line in lines: # 对单行文本进行句子分割 doc = nlp(line) for sent in doc.sents: # 清理句子首尾空白,标准化格式 cleaned_sent = sent.text.strip() all_cleaned_sentences.append(cleaned_sent) # 统计句子出现次数 sent_count = Counter(all_cleaned_sentences) # 转换为DataFrame并按次数降序排序 result_df = pd.DataFrame(sent_count.items(), columns=["Sentence", "Count"]) result_df = result_df.sort_values(by="Count", ascending=False).reset_index(drop=True) # 输出Markdown格式结果 print(result_df.to_markdown(index=False))
代码说明
- CSV读取:批量读取所有目标CSV文件,提取邮件内容列并跳过空值。
- 换行拆分:将每封邮件按换行拆分为独立行,过滤空行避免无效处理。
- 句子分割与标准化:对每行文本单独做句子分割,清理句子首尾空白,保证相同句子格式统一。
- 统计与输出:用
Counter统计句子频次,转换为DataFrame后排序,输出为Markdown表格。
内容的提问来源于stack exchange,提问作者H-Finch
相关产品推荐
相关产品推荐

