You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

requests-html爬取Midjourney渲染失败,无法获取图片求解决方案

解决requests-html爬取Midjourney展示页无法获取图片的问题

问题分析

requests-html的arender()依赖Pyppeteer(无头Chrome),Midjourney页面可能通过检测无头浏览器特征、请求头或渲染环境限制内容加载,导致图片元素无法被捕获;而Selenium使用常规有头浏览器,更接近真实用户环境,因此能正常获取图片。

可行解决方案

核心思路是模拟真实浏览器环境,绕过网站的反爬检测:

  1. 配置真实请求头与视口尺寸
    模拟主流浏览器的User-Agent和屏幕分辨率,避免被识别为无头爬虫。
  2. 调整渲染参数
    延长等待时间确保JS完全执行,添加反检测启动参数,启用图片加载。
  3. 优化元素提取方式
    优先使用CSS选择器或更精确的XPath定位图片元素。

修改后的代码示例

HTMLSession版本

from requests_html import HTMLSession

session = HTMLSession()
# 模拟Chrome浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
response = session.get('https://midjourney.com/showcase/top/', headers=headers)
# 配置渲染参数,绕过无头检测并确保内容加载
response.html.arender(
    timeout=120,
    sleep=10,
    viewport={'width': 1920, 'height': 1080},
    options={
        'args': [
            '--no-sandbox',
            '--disable-setuid-sandbox',
            '--disable-blink-features=AutomationControlled',
            '--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
        ],
        'ignoreDefaultArgs': ['--enable-automation']
    }
)

# 提取图片元素并输出src属性
images = response.html.find('img')
print(f"找到图片数量:{len(images)}")
for img in images:
    print(img.attrs.get('src'))

AsyncHTMLSession版本

from requests_html import AsyncHTMLSession
import asyncio

async def crawl_midjourney_images():
    asession = AsyncHTMLSession()
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    response = await asession.get('https://midjourney.com/showcase/top/', headers=headers)
    await response.html.arender(
        timeout=120,
        sleep=10,
        viewport={'width': 1920, 'height': 1080},
        options={
            'args': [
                '--no-sandbox',
                '--disable-setuid-sandbox',
                '--disable-blink-features=AutomationControlled',
                '--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
            ],
            'ignoreDefaultArgs': ['--enable-automation']
        }
    )
    images = response.html.find('img')
    print(f"找到图片数量:{len(images)}")
    for img in images:
        print(img.attrs.get('src'))

asyncio.run(crawl_midjourney_images())

额外说明

如果上述方案仍无效,说明Midjourney的反爬机制已更新,requests-html的Pyppeteer底层可能无法绕过,此时Selenium仍是更可靠的选择。同时注意控制爬取频率,避免触发网站封禁。

内容的提问来源于stack exchange,提问作者Artin Mohammadi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 11:35:21