You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Requests无法抓取Mangadex网站隐藏数据的问题求助

解决Mangadex最新标题页面抓取时__nuxt标签为空的问题

问题根源

Mangadex采用Nuxt框架构建,核心数据依赖客户端JS动态加载,同时网站存在反爬机制:

  • 默认请求头(尤其是User-Agent)会被识别为爬虫,返回空内容容器
  • requests-html的默认渲染配置未等待页面完全加载,导致JS未完成数据注入

修复方案

1. 优化请求头+延长渲染等待时间

给requests-html添加模拟真实浏览器的请求头,并设置足够的等待时间确保JS执行完毕:

from requests_html import HTMLSession

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5'
}

session = HTMLSession()
response = session.get('https://mangadex.org/titles/latest', headers=headers)
# 等待3秒确保JS加载数据,额外休眠2秒强化稳定性
response.html.render(wait=3, sleep=2)
# 提取__nuxt标签内容
nuxt_content = response.html.find('#__nuxt', first=True).html
print(nuxt_content)

2. 改用Selenium处理人机验证(若触发)

如果上述方法无效,大概率是遇到了Cloudflare人机验证。此时requests-html的无界面浏览器无法通过验证,可改用Selenium配合真实浏览器驱动:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
from bs4 import BeautifulSoup

options = Options()
options.add_argument('--headless=new')
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

driver = webdriver.Chrome(options=options)
driver.get('https://mangadex.org/titles/latest')
# 等待页面加载(若出现手动验证,需关闭headless模式手动完成)
time.sleep(5)
# 解析页面源码
soup = BeautifulSoup(driver.page_source, 'html.parser')
nuxt_content = soup.find('div', id='__nuxt')
print(nuxt_content.prettify())

driver.quit()

3. 优先使用官方API(推荐)

Mangadex提供官方API,直接调用获取数据比爬取页面更稳定合规。例如获取最新发布漫画,可请求对应接口并携带排序、数量参数,返回的JSON数据可直接解析使用,无需处理前端渲染逻辑。

注意事项

  • 控制请求频率,遵守网站使用条款,避免IP被封禁
  • 官方API有请求限制,需合理设置请求间隔

内容的提问来源于stack exchange,提问作者gay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 10:35:06