Python网页爬取:如何在页面加载几秒后获取HTML内容?
解决动态页面爬取问题:获取Toribash论坛最新帖子
嘿,你的问题很典型——requests库只能抓取页面的初始静态HTML,而这个论坛的最新帖子是通过JavaScript动态加载的。你之前加的time.sleep(5)根本没用,因为它是在发送请求前等待,而且requests完全不会执行页面里的JS代码,自然拿不到动态更新后的内容。
下面给你两个可行的解决方案:
方案一:用Selenium模拟浏览器加载动态内容
Selenium可以模拟真实浏览器的行为,执行页面里的JavaScript,等动态内容加载完成后再抓取HTML。这是最直接的方法,适合快速解决问题:
步骤1:安装依赖
先安装Selenium,以及对应浏览器的驱动(以Chrome为例):
pip install selenium
然后下载和你Chrome版本匹配的Chrome浏览器驱动,放到系统可访问的路径里,或者在代码里指定驱动路径。
步骤2:修改你的Discord Bot代码
注意:Discord.py是异步框架,不要用time.sleep()阻塞事件循环,改用Selenium的等待机制:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from lxml import html # 确保你导入了这个库 if message.content.startswith("-post"): await client.send_message(message.channel, ":arrows_counterclockwise: **Accessing forums...**") await client.send_typing(message.channel) # 配置无头浏览器(不弹出可视化窗口) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("--no-sandbox") # 服务器环境运行时可能需要 # 初始化浏览器驱动 driver = webdriver.Chrome(options=chrome_options) url = "http://forum.toribash.com/tori_spy.php" try: driver.get(url) # 等待动态帖子加载完成,最多等10秒 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//strong/a")) ) # 获取加载完成后的完整页面HTML page_html = driver.page_source tree = html.fromstring(page_html) list_stuff = [atag.text_content() for atag in tree.xpath("//strong/a")] if list_stuff: await client.send_message(message.channel, f":white_check_mark: Last post was in the thread **{list_stuff[0]}**") else: await client.send_message(message.channel, ":x: No recent posts found.") except Exception as e: await client.send_message(message.channel, f":x: Failed to load posts: {str(e)}") finally: driver.quit() # 必须关闭浏览器,避免资源泄漏
方案二:直接调用论坛的API(更高效)
如果你想更高效,不想用浏览器模拟,可以打开浏览器开发者工具(F12),切换到Network标签页,刷新tori_spy.php页面,看看有没有XHR/fetch请求是用来获取最新帖子数据的。
这类论坛通常会用API接口动态加载内容,找到对应的接口后,直接用requests请求这个接口,解析返回的JSON数据即可。这种方法速度更快,占用资源更少,不需要依赖浏览器驱动。
举个例子(假设找到API):
if message.content.startswith("-post"): await client.send_message(message.channel, ":arrows_counterclockwise: **Accessing forums...**") await client.send_typing(message.channel) session_requests = requests.session() # 假设找到的API地址是这个(实际需要你自己抓包确认字段和地址) api_url = "http://forum.toribash.com/api/latest_posts" result = session_requests.get(api_url, headers={"referer": "http://forum.toribash.com/tori_spy.php"}) if result.status_code == 200: posts_data = result.json() latest_thread = posts_data[0]["thread_title"] # 字段名需要根据实际返回调整 await client.send_message(message.channel, f":white_check_mark: Last post was in the thread **{latest_thread}**") else: await client.send_message(message.channel, ":x: Failed to fetch posts from API.")
小提示
- 优先尝试方案二,因为API请求比模拟浏览器更稳定高效;
- 如果用Selenium,记得定期更新ChromeDriver,确保和你的Chrome版本匹配;
- 在Discord Bot里,尽量把耗时的操作(比如浏览器加载)放到异步线程里,避免阻塞Bot的其他功能。
内容的提问来源于stack exchange,提问作者johnny mcdangerface
相关产品推荐
相关产品推荐

