Discord机器人开发:如何爬取MMU活动页面「查看更多」后的内容?
问题
我正在用Python开发Discord机器人,需要爬取曼彻斯特城市大学(MMU)活动页面的全部活动数据并发送到Discord频道。目前用requests和bs4实现了初始页面数据爬取,但拿不到点击「查看更多活动」按钮后加载的内容。查看开发者工具网络面板也没找到有效请求,现有代码如下:
from discord.ext import commands import discord import requests from bs4 import BeautifulSoup # Instantiating the bot BOT_TOKEN = 'ABCDE' CHANNEL_ID = 12345 bot = commands.Bot(command_prefix='!', intents=discord.Intents.all()) #event handler that should be called when the bot is ready and connected to Discord. @bot.event async def on_ready(): print("Bot is connected, online and working.") channel = bot.get_channel(CHANNEL_ID) await channel.send("Hello! Work is underway on my project bot!") #command that responds to '!hello' @bot.command() async def hello(ctx): await ctx.send("Hello!") #command that takes in any number of integers and returns the result @bot.command() async def add(ctx, *arr): result = 0 for i in arr: result += int(i) await ctx.send(f"Result: {result}") # This command sends a list of upcoming events at Manchester Metropolitan University to the channel. # It fetches the events page of the MMU website, scrapes the relevant info (Title, date & location) using Beautiful Soup, # and formats it into a string before sending it as a message. @bot.command() async def scrape(ctx): url = 'https://www.mmu.ac.uk/student-life/events/' headers = { 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'Referer': 'https://www.google.com/', 'Connection': 'keep-alive' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') event_list = soup.find('ul', class_='listing--events') event_items = event_list.find_all('li', class_='event-list__item') events = [] for item in event_items: event_name = item.find('h3', class_='event-list__title').text.strip() event_date = item.find('p', class_='event-list__date-start').text.strip() event_location = item.find('p', class_='event-list__location').text.strip() events.append(f"{event_date}\n{event_name}\n{event_location}") await ctx.send("\n\n".join(events)) bot.run(BOT_TOKEN)
解决方案
该页面的「查看更多活动」属于前端动态加载内容,并非通过独立API请求返回数据,所以直接用requests无法获取后续内容。可以用playwright模拟真实浏览器行为来加载全部活动数据:
步骤1:安装依赖
先安装playwright并配置Chromium浏览器:
pip install playwright playwright install chromium
步骤2:修改scrape命令代码
替换原有的scrape函数,用playwright模拟点击加载更多,直到所有内容加载完成:
from discord.ext import commands import discord from bs4 import BeautifulSoup from playwright.async_api import async_playwright # Instantiating the bot BOT_TOKEN = 'ABCDE' CHANNEL_ID = 12345 bot = commands.Bot(command_prefix='!', intents=discord.Intents.all()) #event handler that should be called when the bot is ready and connected to Discord. @bot.event async def on_ready(): print("Bot is connected, online and working.") channel = bot.get_channel(CHANNEL_ID) await channel.send("Hello! Work is underway on my project bot!") #command that responds to '!hello' @bot.command() async def hello(ctx): await ctx.send("Hello!") #command that takes in any number of integers and returns the result @bot.command() async def add(ctx, *arr): result = 0 for i in arr: result += int(i) await ctx.send(f"Result: {result}") # This command sends a list of upcoming events at Manchester Metropolitan University to the channel. # It uses playwright to simulate browser behavior, load all events by clicking "View more events", # then scrapes the relevant info using Beautiful Soup and sends it as a message. @bot.command() async def scrape(ctx): url = 'https://www.mmu.ac.uk/student-life/events/' async with async_playwright() as p: # 启动无头Chromium浏览器 browser = await p.chromium.launch(headless=True) page = await browser.new_page( user_agent='Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/110.0.0.0 Safari/537.36' ) await page.goto(url) # 循环点击"查看更多活动"按钮,直到按钮消失 while True: try: # 等待按钮出现并点击 await page.click('button:has-text("View more events")', timeout=3000) # 等待页面加载新内容 await page.wait_for_timeout(1000) except: # 按钮不存在时退出循环 break # 获取完整页面HTML page_html = await page.content() await browser.close() # 解析HTML soup = BeautifulSoup(page_html, 'html.parser') event_list = soup.find('ul', class_='listing--events') event_items = event_list.find_all('li', class_='event-list__item') events = [] for item in event_items: event_name = item.find('h3', class_='event-list__title').text.strip() event_date = item.find('p', class_='event-list__date-start').text.strip() event_location = item.find('p', class_='event-list__location').text.strip() events.append(f"{event_date}\n{event_name}\n{event_location}") # 发送结果,Discord单条消息有字符限制,若内容过长需拆分 if len("\n\n".join(events)) > 2000: # 拆分消息 chunk = [] current_length = 0 for event in events: event_length = len(event) + 2 # 加上分隔符长度 if current_length + event_length > 2000: await ctx.send("\n\n".join(chunk)) chunk = [event] current_length = event_length else: chunk.append(event) current_length += event_length if chunk: await ctx.send("\n\n".join(chunk)) else: await ctx.send("\n\n".join(events)) bot.run(BOT_TOKEN)
关键说明
playwright会模拟真实浏览器加载页面,包括执行JavaScript代码,因此能获取到点击「查看更多」后加载的所有内容- 循环点击按钮直到无法找到,确保所有活动都被加载
- 增加了消息拆分逻辑,避免因内容过长超过Discord单条消息的字符限制(2000字符)
内容的提问来源于stack exchange,提问作者Jen L
相关产品推荐
相关产品推荐

