requests-html爬取Midjourney渲染失败,无法获取图片求解决方案
解决requests-html爬取Midjourney展示页无法获取图片的问题
问题分析
requests-html的arender()依赖Pyppeteer(无头Chrome),Midjourney页面可能通过检测无头浏览器特征、请求头或渲染环境限制内容加载,导致图片元素无法被捕获;而Selenium使用常规有头浏览器,更接近真实用户环境,因此能正常获取图片。
可行解决方案
核心思路是模拟真实浏览器环境,绕过网站的反爬检测:
- 配置真实请求头与视口尺寸
模拟主流浏览器的User-Agent和屏幕分辨率,避免被识别为无头爬虫。 - 调整渲染参数
延长等待时间确保JS完全执行,添加反检测启动参数,启用图片加载。 - 优化元素提取方式
优先使用CSS选择器或更精确的XPath定位图片元素。
修改后的代码示例
HTMLSession版本
from requests_html import HTMLSession session = HTMLSession() # 模拟Chrome浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = session.get('https://midjourney.com/showcase/top/', headers=headers) # 配置渲染参数,绕过无头检测并确保内容加载 response.html.arender( timeout=120, sleep=10, viewport={'width': 1920, 'height': 1080}, options={ 'args': [ '--no-sandbox', '--disable-setuid-sandbox', '--disable-blink-features=AutomationControlled', '--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' ], 'ignoreDefaultArgs': ['--enable-automation'] } ) # 提取图片元素并输出src属性 images = response.html.find('img') print(f"找到图片数量:{len(images)}") for img in images: print(img.attrs.get('src'))
AsyncHTMLSession版本
from requests_html import AsyncHTMLSession import asyncio async def crawl_midjourney_images(): asession = AsyncHTMLSession() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = await asession.get('https://midjourney.com/showcase/top/', headers=headers) await response.html.arender( timeout=120, sleep=10, viewport={'width': 1920, 'height': 1080}, options={ 'args': [ '--no-sandbox', '--disable-setuid-sandbox', '--disable-blink-features=AutomationControlled', '--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' ], 'ignoreDefaultArgs': ['--enable-automation'] } ) images = response.html.find('img') print(f"找到图片数量:{len(images)}") for img in images: print(img.attrs.get('src')) asyncio.run(crawl_midjourney_images())
额外说明
如果上述方案仍无效,说明Midjourney的反爬机制已更新,requests-html的Pyppeteer底层可能无法绕过,此时Selenium仍是更可靠的选择。同时注意控制爬取频率,避免触发网站封禁。
内容的提问来源于stack exchange,提问作者Artin Mohammadi
相关产品推荐
相关产品推荐

