如何用Curl绕过nhentai的Cloudflare JS挑战获取画廊封面图
解决方案:nhentai画廊封面图获取(绕过403与Cloudflare验证)
一、解决Cloudflare拦截与curl 403问题
Cloudflare的拦截核心是检测请求是否来自真实浏览器,以下是轻量解决方法:
- 添加完整浏览器请求头:直接curl的默认请求头会被识别为非浏览器请求,需要模拟真实浏览器的请求参数:
curl -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \ -H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8" \ -H "Accept-Language: en-US,en;q=0.5" \ -H "Connection: keep-alive" \ -H "Upgrade-Insecure-Requests: 1" \ https://nhentai.net/g/177013/ - 持久化会话Cookie:第一次请求页面获取Cloudflare验证后的Cookie,后续请求复用可避免重复验证:
# 首次请求存储Cookie curl -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \ --cookie-jar nhentai_cookies.txt \ https://nhentai.net/g/177013/ # 复用Cookie请求页面 curl -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \ --cookie nhentai_cookies.txt \ https://nhentai.net/g/177013/ - 若遇到人机验证:需手动用浏览器打开画廊页面完成验证,然后导出Cookie到本地文件,再用
--cookie参数复用。
二、获取正确的封面图路径(解决ID与图片服务器不匹配问题)
无法通过画廊ID直接推导图片路径,必须从画廊页面源码中提取真实URL:
- Shell命令一键提取下载(适用于Linux/macOS):
# 提取封面URL COVER_URL=$(curl -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \ --cookie nhentai_cookies.txt \ https://nhentai.net/g/177013/ | grep -o 'https://t[0-9]\.nhentai\.net/galleries/[0-9]*/cover\.[a-z]*' | head -1) # 下载封面 curl -A "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" \ --cookie nhentai_cookies.txt \ -o cover_177013.jpg \ $COVER_URL - 轻量Python脚本实现(无重型依赖,仅用内置库):
import urllib.request import re # 配置请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } gallery_id = '177013' gallery_url = f'https://nhentai.net/g/{gallery_id}/' # 获取画廊页面源码 req = urllib.request.Request(gallery_url, headers=headers) with urllib.request.urlopen(req) as response: html_content = response.read().decode('utf-8') # 正则提取封面URL cover_pattern = re.compile(r'https://t\d\.nhentai\.net/galleries/\d+/cover\.[a-z]+') cover_url = cover_pattern.search(html_content).group() # 下载封面 img_req = urllib.request.Request(cover_url, headers=headers) with urllib.request.urlopen(img_req) as img_response: with open(f'cover_{gallery_id}.jpg', 'wb') as img_file: img_file.write(img_response.read()) print(f"封面已保存为:cover_{gallery_id}.jpg")
三、常见问题排查
- 若仍返回Cloudflare验证页面:检查User-Agent是否为最新浏览器版本,或确保Cookie是刚验证过的有效Cookie。
- 提取封面URL失败:可能是页面结构更新,需调整正则表达式,比如匹配
<img class="cover" src="..."中的src属性值。
内容的提问来源于stack exchange,提问作者DavidNyan10
相关产品推荐
相关产品推荐

