You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫求助:Requests-html无法获取漫画网站页面内容

解决漫画网站爬取时JS Cookie验证导致无法获取页面内容的问题

我是Python新手,近期在做爬取漫画网站章节图片的小项目时遇到问题:使用BeautifulSoup.select()无法得到预期结果,测试打印页面内容时,仅返回一段设置Cookie并刷新页面的JS代码:

'document.cookie="VinaHost-Shield=a7a00919549a80aa44d5e1df8a26ae20"+"; path=/";window.location.reload(true);'

尝试过的无效代码

方法1:使用requests_html渲染页面

from requests_html import HTMLSession
session = HTMLSession()

res = session.get("https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html")
res.html.render()
print(res.content)

方法2:直接使用requests+BeautifulSoup

import requests, bs4

url = "https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html"
res = requests.get(url, headers={"User-Agent": "Requests"})
res.raise_for_status()
# soup = bs4.BeautifulSoup(res.text, "html.parser")
# onePiece = soup.select(".page-chapter")
print(res.content)

解决方法及最终可用代码

在Windows 11上安装Docker和Splash后问题得到解决,最终实现代码如下:

import os
import requests, bs4
os.makedirs("OnePiece", exist_ok=True)
url = "https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html"
res = requests.get("http://localhost:8050/render.html", params={"url": url, "wait": 5})
res.raise_for_status()
soup = bs4.BeautifulSoup(res.text, "html.parser")
onePiece = soup.find_all("img", class_="lazy")
for element in onePiece:
    imageLink = "https:" + element["data-cdn"]
    res = requests.get(imageLink)
    imageFile = open(os.path.join("OnePiece", os.path.basename(imageLink)), "wb")
    for chunk in res.iter_content(100000):
        imageFile.write(chunk)
    imageFile.close()

内容的提问来源于stack exchange,提问作者Jim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 02:25:40