如何在Python中用Counter过滤WhatsApp聊天常用词并排除指定词汇
解决WhatsApp聊天记录词汇统计时排除特定内容的问题
作为Python新手,你已经迈出了很好的第一步!要忽略日期、自己和对方的名字,只需要在统计前添加过滤步骤即可,这里给你两种实用的方案:
方案一:精确排除指定词汇
这种方法适合你明确知道要忽略的具体名字和日期的情况,直接把它们放进排除列表就行:
from collections import Counter import re # 替换成你实际要忽略的名字、日期等词汇,用集合查询效率更高 exclude_words = {"你的名字", "对方的名字", "10/05/2024", "15-06-2024"} # 用with语句打开文件更安全,避免资源泄漏 with open('chat.txt', 'r') as chat_file: chat_content = chat_file.read().lower() # 提取所有单词 all_words = re.findall(r'\w+', chat_content) # 过滤掉排除列表里的词汇 filtered_words = [word for word in all_words if word not in exclude_words] # 输出最常用的10个词汇 print(Counter(filtered_words).most_common(10))
方案二:自动排除日期格式的内容
如果聊天里的日期格式比较固定(比如dd/mm/yyyy或dd-mm-yyyy),可以用正则表达式自动识别并排除,再加上名字过滤:
from collections import Counter import re # 替换成你要忽略的名字 exclude_names = {"你的名字", "对方的名字"} with open('chat.txt', 'r') as chat_file: chat_content = chat_file.read().lower() all_words = re.findall(r'\w+', chat_content) # 双重过滤:既不是要排除的名字,也不是日期格式 filtered_words = [ word for word in all_words if word not in exclude_names and not re.match(r'^\d{1,2}[-/]\d{1,2}[-/]\d{2,4}$', word) ] print(Counter(filtered_words).most_common(10))
额外小贴士
- 如果要排除的词汇很多,可以把它们写在一个
exclude.txt文件里,每行一个词汇,然后用exclude_words = set(line.strip().lower() for line in open('exclude.txt'))读取,更方便维护。 - 要是聊天里还有其他不想统计的系统内容(比如WhatsApp的
media omitted),直接把对应的词汇加入排除列表就行。
内容的提问来源于stack exchange,提问作者PEOlhc
相关产品推荐
相关产品推荐

