Telegram Bot如何在触发MessageHandler前过滤链接避免正则误匹配
解决方案
以下三种方案均可实现需求,推荐优先选择方案1,改动最小且逻辑最清晰:
方案1:将匹配逻辑移入回调函数内部预处理
直接调整Handler的触发规则,将文本清洗和规则匹配放到业务函数中执行,不需要修改Handler触发流程,适配性最强:
- 修改
entry_points的过滤规则为纯文本匹配 - 在业务函数开头先移除文本中的所有链接,再做正则匹配,不匹配直接终止流程即可
修改后的完整代码如下:
import re import random from telegram import Update from telegram.ext import ConversationHandler, MessageHandler, Filters, CallbackContext class EveryOtherQuestion(ConversationHandler): def __init__(self): super().__init__( # 改为匹配所有文本消息 entry_points=[MessageHandler(Filters.text, EveryOtherQuestion.everyOtherQuestion)], states={}, fallbacks=[] ) @staticmethod def everyOtherQuestion(update: Update, context: CallbackContext): raw_text = update.effective_message.text or '' # 先移除所有链接内容 cleaned_text = re.sub(r'http\S+', '', raw_text).strip() # 再执行问句匹配 if not re.search(r".{2,}\?+", cleaned_text, re.IGNORECASE): return ConversationHandler.END answers = [ "idT", "Gwiss ned!", "Auf keinen Fall", "Da kannst drauf wetten", "Natürlich nicht! So ein Bach!", "Ja", "Nein", "Wie kommst du denn da drauf?! Spinnst gwiss a weng?!", ] update.message.reply_text(text=random.choice(answers), quote=True) return ConversationHandler.END
方案2:调整正则表达式,直接排除链接内的问号
如果不想改动现有Handler结构,可以通过负向预查修改正则规则,直接忽略URL中包含的问号:
将原来的正则改为:re.compile(r"(?<!http\S).{2,}\?+", re.IGNORECASE)
该规则的作用是匹配问号前不是URL字符的问句,不会命中链接里的问号。
注意:如果文本同时包含普通问句和带问号的链接,该规则也会正常触发,符合需求
方案3:添加全局预处理Handler
如果有多处业务都需要用到清洗后的文本,可以在所有Handler注册前,先注册一个全局预处理Handler,优先级最高,执行文本清洗后将结果存入context供后续Handler调用:
def preprocess_handler(update: Update, context: CallbackContext): if update.effective_message.text: # 全局统一清洗文本 cleaned_text = re.sub(r'http\S+', '', update.effective_message.text).strip() context.bot_data['cleaned_text'] = cleaned_text # 无返回值,Handler会继续向下匹配
注册时需要将该Handler放在最前面:
application.add_handler(MessageHandler(Filters.text, preprocess_handler)) application.add_handler(EveryOtherQuestion())
后续所有业务Handler中都可以直接取context.bot_data['cleaned_text']使用。
内容的提问来源于stack exchange,提问作者Jona
相关产品推荐
相关产品推荐

