You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium/BeautifulSoup抓取jQuery渲染的动漫网站表格

解决方案:用Selenium等待渲染后抓取HTML就行

你不用纠结找XHR或者那个jQuery文件,这个网站的排行榜内容是前端渲染的,直接用Selenium等页面加载完抓渲染后的HTML就可以搞定,步骤和代码参考下面:

  • 步骤1:配置Selenium并等待页面渲染
    启动浏览器后,不要立刻拿源码,用WebDriverWait等待排行榜的列表元素加载出来,确保内容已经渲染完成。比如等第一个漫画项出现,或者整个排行榜容器加载完毕。

  • 步骤2:提取渲染后的HTML并解析
    拿到页面源码后,用BeautifulSoup解析定位元素,提取你要的前十项内容。

示例代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 初始化浏览器(需对应版本的ChromeDriver)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")  # 无头模式,可选
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
driver = webdriver.Chrome(options=options)

try:
    driver.get("https://www.anime-planet.com/manga/top-manga/week")
    # 等待排行榜列表加载,可根据实际元素调整定位器
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, ".tableList tbody tr"))
    )
    # 获取渲染后的页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    
    # 提取前十项漫画
    top_manga = soup.select(".tableList tbody tr")[:10]
    for idx, item in enumerate(top_manga, 1):
        title = item.find("a", class_="title").text.strip()
        rating = item.find("td", class_="numeric").text.strip()
        print(f"第{idx}名:{title},评分:{rating}")
finally:
    driver.quit()

补充说明:

  • 为什么不用找XHR?这个网站的数据不是通过单独的XHR接口返回的,而是前端JS直接把数据嵌在页面里或者动态生成DOM,所以抓渲染后的HTML是最直接的方式。
  • 注意反爬:不要频繁请求,可加随机等待时间;如果遇到验证码或者IP封禁,可能需要降低爬取频率。
  • 如果不想用Selenium,也可以尝试搜索页面里的脚本标签,看有没有内嵌的JSON数据(比如window.开头的变量),解析JSON也能拿到数据,但这种方式依赖网站代码结构,不如Selenium通用。

内容的提问来源于stack exchange,提问作者moosepowa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 03:30:50