Python爬虫求助:Requests-html无法获取漫画网站页面内容
我是Python新手,近期在做爬取漫画网站章节图片的小项目时遇到问题:使用BeautifulSoup.select()无法得到预期结果,测试打印页面内容时,仅返回一段设置Cookie并刷新页面的JS代码:
'document.cookie="VinaHost-Shield=a7a00919549a80aa44d5e1df8a26ae20"+"; path=/";window.location.reload(true);'
尝试过的无效代码
方法1:使用requests_html渲染页面
from requests_html import HTMLSession session = HTMLSession() res = session.get("https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html") res.html.render() print(res.content)
方法2:直接使用requests+BeautifulSoup
import requests, bs4 url = "https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html" res = requests.get(url, headers={"User-Agent": "Requests"}) res.raise_for_status() # soup = bs4.BeautifulSoup(res.text, "html.parser") # onePiece = soup.select(".page-chapter") print(res.content)
解决方法及最终可用代码
在Windows 11上安装Docker和Splash后问题得到解决,最终实现代码如下:
import os import requests, bs4 os.makedirs("OnePiece", exist_ok=True) url = "https://truyenqqpro.com/truyen-tranh/dao-hai-tac-128-chap-1060.html" res = requests.get("http://localhost:8050/render.html", params={"url": url, "wait": 5}) res.raise_for_status() soup = bs4.BeautifulSoup(res.text, "html.parser") onePiece = soup.find_all("img", class_="lazy") for element in onePiece: imageLink = "https:" + element["data-cdn"] res = requests.get(imageLink) imageFile = open(os.path.join("OnePiece", os.path.basename(imageLink)), "wb") for chunk in res.iter_content(100000): imageFile.write(chunk) imageFile.close()
内容的提问来源于stack exchange,提问作者Jim
相关产品推荐
相关产品推荐

