如何使用Python正则移除Discord聊天TXT文件中的发件人名称与@标签
实现方案
调整后的完整可运行代码如下:
import collections import pandas as pd import matplotlib.pyplot as plt import re # 新增导入正则处理模块 # 1. 读取并清理Discord对话内容 cleaned_msgs = [] with open('generic_discord_talk.txt', encoding="utf8") as talk_f: for line in talk_f: # 过滤无冒号的无效行,拆分丢弃冒号前的发件人信息 if ': ' not in line: continue msg_content = line.split(': ', 1)[1] # 全局替换所有<@!数字ID>格式的@标签为空 msg_clean = re.sub(r'<@!\d+>', '', msg_content).strip() cleaned_msgs.append(msg_clean) # 2. 读取停用词表,合并自定义停用词 with open('stopwords-es.txt', encoding="utf8") as stop_f: es_stops = [word.strip() for line in stop_f for word in line.split()] stopwords = set(es_stops).union(set(['you','for','the'])) # 3. 拆分清理后的对话为单词,过滤停用词得到最终列表 all_words = [] for msg in cleaned_msgs: # 拆分单词同时去除首尾标点 words = [word.strip('.,?!') for word in msg.split() if word.strip('.,?!')] all_words.extend(words) st = [word for word in all_words if word.lower() not in stopwords] print(st)
核心逻辑说明
- 用
split(': ', 1)仅拆分每行第一个冒号,直接丢弃冒号前所有发件人内容,无需提前匹配具体发件人名称 - 正则
r'<@!\d+>'可精准匹配所有格式为<@!数字ID>的提及标签,re.sub会全局替换所有匹配内容,无需仅匹配行首 - 分词时额外清理了单词首尾的标点符号,避免标点和单词绑定影响停用词过滤效果
内容的提问来源于stack exchange,提问作者user15898019
相关产品推荐
相关产品推荐

