如何在Discord.py中运行Selenium爬虫循环且不阻塞命令?
Discord机器人爬虫阻塞命令响应问题
问题描述
我用discord.py做了个测试用Discord机器人,内置基础命令。还加了个每X秒运行的定时任务,循环调用另一个用Selenium+BeautifulSoup写的推特搜索爬虫。现在的问题是:爬虫运行时,机器人所有命令都用不了,只有爬虫结束后才恢复正常——哪怕爬虫在sleep等页面加载的时候,命令还是被阻塞。
我试过用threading开新线程跑爬虫,但没弄对,问题依旧。刚接触Python,不知道怎么推进。
相关代码
discordbot.py
import discord import bs4 from bs4 import BeautifulSoup from discord.ext import commands from discord.ext import tasks from discord import app_commands import threading import time from time import sleep from twitterscrapper import scrapeSearch searchInterval = 30 searchChannelID = 0000000 BOT_TOKEN = "TOKEN" CHANNEL_ID = 000000 intents = discord.Intents.default() intents.presences = True intents.members = True client = discord.Client(intents=intents) tree = app_commands.CommandTree(client) bot = commands.Bot(command_prefix="!", intents=intents) @bot.event async def on_ready(): print("Bot Online") channel = bot.get_channel(CHANNEL_ID) await channel.send("Hello! Bot is online") try: synced = await bot.tree.sync() # globally sync slash commands print(f"Synced {len(synced)} commands") except Exception as e: print(e) scrapingLoop.start() @bot.tree.command(name="ping", description="pong") async def ping(interaction: discord.Interaction): await interaction.response.send_message("Pong!", ephemeral=True) @bot.tree.command(name="memberinfo", description="Displays User Info") async def memberinfo(interaction: discord.Interaction, member: discord.Member = None): if member == None: member = interaction.user embed = discord.Embed(title="Member Info") embed.add_field(name="ID", value=member.id) embed.add_field(name="Name", value=member.name) embed.add_field(name="Status", value=interaction.guild.get_member(member.id).status) embed.add_field(name="Creation Date", value=member.created_at.strftime("%m/%d/%Y")) await interaction.response.send_message(embed=embed) @tasks.loop(seconds=searchInterval) async def scrapingLoop(): print("running loop") link = scrapeSearch() if link != "No new posts": searchChannel = bot.get_channel(searchChannelID) await searchChannel.send(str(link)) else: searchChannel = bot.get_channel(searchChannelID) await searchChannel.send( "No new posts found in the last " + str(searchInterval) + " seconds" ) bot.run(BOT_TOKEN)
twitterscraper.py
from bs4 import BeautifulSoup from selenium import webdriver from urllib.parse import unquote from selenium import webdriver from selenium.webdriver import ChromeOptions, Keys from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.wait import WebDriverWait import time from time import sleep target_url = "TWITTER SEARCH URL" def scrapeSearch(): options = ChromeOptions() options.add_argument("--start-maximized") options.add_experimental_option("excludeSwitches", ["enable-automation"]) driver = webdriver.Chrome(options=options) driver.get(target_url) username = WebDriverWait(driver, 20).until( EC.visibility_of_element_located( (By.CSS_SELECTOR, 'input[autocomplete="username"]') ) ) username.send_keys("PHONE_NUMBER") username.send_keys(Keys.ENTER) password = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CSS_SELECTOR, 'input[name="password"]')) ) password.send_keys("PASSWORD") password.send_keys(Keys.ENTER) time.sleep(4) resp = driver.page_source soup = BeautifulSoup(resp, "html.parser") first_result = soup.find("article", {"data-testid": "tweet"}) first_result_anchor_children = first_result.findChildren(name="a") for child in first_result_anchor_children: link = child.get("href") if "status" in link: link = "https://twitter.com" + link f = open("recentlink.txt", "r") lastLink = f.read() if lastLink == link: driver.close() print("1") link = "No new posts" return link else: f = open("recentlink.txt", "w") f.write(link) print("2") driver.close() return link break print(link)
解决方案
问题根源是爬虫是同步阻塞代码,在discord.py的事件循环线程里运行时,会卡住整个线程,导致机器人无法处理命令。以下是两种可行的解决方式:
方法1:用asyncio.run_in_executor(推荐,贴合discord.py异步架构)
修改scrapingLoop函数,把同步的爬虫任务放到线程池执行,避免阻塞事件循环:
import asyncio # 需导入asyncio模块 @tasks.loop(seconds=searchInterval) async def scrapingLoop(): print("running loop") # 用线程池执行同步爬虫函数 loop = asyncio.get_event_loop() link = await loop.run_in_executor(None, scrapeSearch) if link != "No new posts": searchChannel = bot.get_channel(searchChannelID) await searchChannel.send(str(link)) else: searchChannel = bot.get_channel(searchChannelID) await searchChannel.send( f"No new posts found in the last {searchInterval} seconds" )
方法2:用threading手动开线程(需处理异步回调)
如果坚持用threading,注意不能直接在线程里调用discord的异步API,要把结果传回事件循环处理:
from functools import partial def run_scrape(callback): link = scrapeSearch() # 将结果回调到事件循环处理 asyncio.run_coroutine_threadsafe(callback(link), bot.loop) async def handle_scrape_result(link): if link != "No new posts": searchChannel = bot.get_channel(searchChannelID) await searchChannel.send(str(link)) else: searchChannel = bot.get_channel(searchChannelID) await searchChannel.send( f"No new posts found in the last {searchInterval} seconds" ) @tasks.loop(seconds=searchInterval) async def scrapingLoop(): print("running loop") # 开新线程执行爬虫,传入结果处理函数 threading.Thread(target=run_scrape, args=(partial(handle_scrape_result),)).start()
额外优化建议
- 复用浏览器实例:每次爬虫都新开Chrome实例效率极低,可将driver初始化移到函数外复用,避免重复启动关闭。
- 文件操作规范:读写
recentlink.txt时用with语句自动管理文件资源:with open("recentlink.txt", "r") as f: lastLink = f.read() # 写入同理 with open("recentlink.txt", "w") as f: f.write(link) - 增加错误处理:给爬虫代码加try-except块,避免单次爬取出错导致定时任务崩溃:
def scrapeSearch(): try: # 原爬虫代码 except Exception as e: print(f"爬虫出错: {e}") if 'driver' in locals(): driver.close() return "No new posts"
内容的提问来源于stack exchange,提问作者kingo
相关产品推荐
相关产品推荐

