You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python正则移除Discord聊天TXT文件中的发件人名称与@标签

实现方案

调整后的完整可运行代码如下:

import collections
import pandas as pd
import matplotlib.pyplot as plt
import re  # 新增导入正则处理模块

# 1. 读取并清理Discord对话内容
cleaned_msgs = []
with open('generic_discord_talk.txt', encoding="utf8") as talk_f:
    for line in talk_f:
        # 过滤无冒号的无效行,拆分丢弃冒号前的发件人信息
        if ': ' not in line:
            continue
        msg_content = line.split(': ', 1)[1]
        # 全局替换所有<@!数字ID>格式的@标签为空
        msg_clean = re.sub(r'<@!\d+>', '', msg_content).strip()
        cleaned_msgs.append(msg_clean)

# 2. 读取停用词表,合并自定义停用词
with open('stopwords-es.txt', encoding="utf8") as stop_f:
    es_stops = [word.strip() for line in stop_f for word in line.split()]
stopwords = set(es_stops).union(set(['you','for','the']))

# 3. 拆分清理后的对话为单词,过滤停用词得到最终列表
all_words = []
for msg in cleaned_msgs:
    # 拆分单词同时去除首尾标点
    words = [word.strip('.,?!') for word in msg.split() if word.strip('.,?!')]
    all_words.extend(words)
st = [word for word in all_words if word.lower() not in stopwords]

print(st)

核心逻辑说明

  • 用split(': ', 1)仅拆分每行第一个冒号,直接丢弃冒号前所有发件人内容,无需提前匹配具体发件人名称
  • 正则r'<@!\d+>'可精准匹配所有格式为<@!数字ID>的提及标签,re.sub会全局替换所有匹配内容,无需仅匹配行首
  • 分词时额外清理了单词首尾的标点符号,避免标点和单词绑定影响停用词过滤效果

内容的提问来源于stack exchange,提问作者user15898019

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 13:45:02