You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Telegram爬虫Bot终止功能求助:点击按钮停止parsing脚本

解决Telegram爬虫Bot终止爬取任务的问题

核心结论:必须修改你的parsing()脚本,让它支持可中断逻辑,同时在Bot代码中配合管理中断信号——因为Telegram Bot的事件(按钮点击)是异步响应的,若parsing()是阻塞式的长任务,Bot主线程会被卡住无法处理按钮事件;就算是异步任务,也需要在爬取逻辑中加入中断检查点。

为什么之前的方法没生效?

  • 全局变量:多用户场景下会互相干扰,且如果parsing()是同步阻塞任务,Bot主线程被卡住根本没法修改全局变量。
  • FSMContext:它仅用于管理对话状态流转,无法主动终止正在运行的后台任务,只能记录状态,没法中断执行中的函数。

具体实现方案

1. 重构parsing()为可中断任务

根据你的爬取逻辑类型,选择对应的中断方式:

异步爬取场景(推荐用aiohttp替代requests)

import asyncio
from aiohttp import ClientSession

async def parsing_task(stop_event: asyncio.Event, user_id: int):
    async with ClientSession() as session:
        # 模拟分页爬取循环
        for page_num in range(1, 100):
            # 关键:每次迭代前检查终止信号
            if stop_event.is_set():
                print(f"用户{user_id}已终止爬取")
                break
            
            # 实际爬取逻辑
            resp = await session.get(f"https://target-site.com/page/{page_num}")
            page_content = await resp.text()
            # 处理页面内容(解析、存储等)
            await asyncio.sleep(0.5)  # 模拟处理延迟

同步爬取场景(无法改成异步时)

用线程池运行同步任务,配合threading.Event做中断:

import threading
import requests

def parsing_task(stop_event: threading.Event, user_id: int):
    for page_num in range(1, 100):
        if stop_event.is_set():
            print(f"用户{user_id}已终止爬取")
            break
        
        resp = requests.get(f"https://target-site.com/page/{page_num}")
        page_content = resp.text
        # 处理内容...

2. 在Bot代码中管理中断信号

给每个用户维护独立的中断事件,避免多用户冲突:

from telegram import Update, InlineKeyboardButton, InlineKeyboardMarkup
from telegram.ext import ApplicationBuilder, CommandHandler, CallbackQueryHandler, ContextTypes
import asyncio
import threading

# 存储每个用户的中断事件(key: 用户ID, value: 中断事件对象)
user_stop_signals = {}

# 启动爬取的命令处理
async def start_parse(update: Update, context: ContextTypes.DEFAULT_TYPE):
    user_id = update.effective_user.id
    # 创建对应用户的中断事件
    if user_id in user_stop_signals:
        await update.message.reply_text("你已有正在运行的爬取任务")
        return
    
    # 根据爬取类型选择asyncio.Event或threading.Event
    stop_event = asyncio.Event()  # 异步爬取用这个
    # stop_event = threading.Event()  # 同步爬取用这个
    user_stop_signals[user_id] = stop_event

    # 发送带终止按钮的消息
    keyboard = [[InlineKeyboardButton("Finish Parsing", callback_data="stop_parse")]]
    await update.message.reply_text(
        "爬取已启动,点击下方按钮终止",
        reply_markup=InlineKeyboardMarkup(keyboard)
    )

    # 启动爬取任务
    if isinstance(stop_event, asyncio.Event):
        # 异步任务直接用create_task
        asyncio.create_task(parsing_task(stop_event, user_id))
    else:
        # 同步任务放到线程池运行
        threading.Thread(target=parsing_task, args=(stop_event, user_id), daemon=True).start()

# 终止按钮的回调处理
async def stop_parse_callback(update: Update, context: ContextTypes.DEFAULT_TYPE):
    query = update.callback_query
    await query.answer()
    user_id = query.from_user.id

    if user_id not in user_stop_signals:
        await query.edit_message_text("没有正在运行的爬取任务")
        return
    
    # 触发终止信号
    user_stop_signals[user_id].set()
    await query.edit_message_text("爬取已终止")
    # 清理资源,避免内存泄漏
    del user_stop_signals[user_id]

if __name__ == "__main__":
    app = ApplicationBuilder().token("你的BotToken").build()
    app.add_handler(CommandHandler("start", start_parse))
    app.add_handler(CallbackQueryHandler(stop_parse_callback, pattern="^stop_parse$"))
    app.run_polling()

关键注意事项

  • 一定要给每个用户单独维护中断事件,不能用全局变量,否则多用户同时使用时会互相干扰。
  • 任务终止或完成后,记得删除对应的中断事件对象,避免内存泄漏。
  • 同步爬取任务要放到线程/进程池运行,不能直接在Bot主线程执行,否则会卡住Bot无法响应其他事件。

内容的提问来源于stack exchange,提问作者KompyterWik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 23:33:23