You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests_html提取JS内容失败的问题及解决方案咨询

爬虫动态内容提取问题解答

问题背景

开发网络爬虫时,用Beautiful Soup处理静态HTML正常,用requests_html提取动态JS内容在多数网站有效,但爬取https://goglobal.com/时,无法获取“100+ countries”“2500+ employees”“3 Billion dollars saved...”区域的数值,测试脚本已尝试延长等待时间但无效。

测试脚本:

from requests_html import HTMLSession
import time
session = HTMLSession()
url = "https://goglobal.com/"
r = session.get(url)

r.html.render(wait=10)
time.sleep(10)
print(r.html.html)

1. 为何内容无法正确加载?

  • 反爬检测:网站可能识别出requests_html依赖的无头Pyppeteer浏览器特征,限制了动态内容渲染,这类反爬会针对navigator.webdriver等属性做拦截。
  • 触发条件未满足:这些统计数字需要页面滚动到对应区域才会触发JS渲染逻辑,requests_html默认只渲染首屏内容,无滚动动作则不会加载目标内容。
  • JS环境缺失:requests_html的Pyppeteer环境可能缺少网站所需的部分JS API或特性,导致统计数字的渲染代码执行失败。

2. 能否继续使用requests_html解决此问题?

可以尝试以下调整方案,仍有概率解决问题:

  • 添加真实User-Agent:模拟普通浏览器请求,避免被识别为爬虫
  • 设置滚动触发:通过scrolldown参数让页面滚动,触发内容加载
  • 调整等待逻辑:延长渲染等待时间,确保JS完全执行

调整后的测试脚本:

from requests_html import HTMLSession
session = HTMLSession()
url = "https://goglobal.com/"
# 自定义浏览器User-Agent
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}
r = session.get(url, headers=headers)
# 滚动2次触发内容加载,每次间隔2秒
r.html.render(wait=5, scrolldown=2, sleep=2)
# 提取目标元素
counters = r.html.find('.counter-number')
for counter in counters:
    print(counter.text)

如果上述调整仍无效,说明网站反爬机制针对Pyppeteer做了强拦截,requests_html的局限性会凸显,难以绕过。

3. 是否可通过selenium、playwright或scrapy解决此问题?

这三个工具都可以解决,具体方案如下:

  • Playwright:对无头浏览器的伪装性最好,内置反爬规避特性,代码实现简洁:
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
    page.goto("https://goglobal.com/", wait_until="networkidle")
    # 滚动到统计区域触发渲染
    page.locator('.counter-number').first.scroll_into_view_if_needed()
    page.wait_for_selector('.counter-number', state='visible')
    counters = page.locator('.counter-number').all_text_contents()
    print(counters)
    browser.close()
  • Selenium:兼容性强,可通过修改浏览器参数规避自动化检测:
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = Options()
options.add_argument('--headless=new')
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
# 禁用自动化检测特征
options.add_argument('--disable-blink-features=AutomationControlled')

driver = webdriver.Chrome(options=options)
driver.get("https://goglobal.com/")
# 等待元素加载并滚动到目标区域
wait = WebDriverWait(driver, 15)
counters = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'counter-number')))
for counter in counters:
    driver.execute_script("arguments[0].scrollIntoView();", counter)
    print(counter.text)
driver.quit()
  • Scrapy:本身是静态爬虫框架,但可通过scrapy-playwright中间件集成Playwright处理动态内容,适合大规模爬取场景:
  1. 安装依赖:pip install scrapy-playwright
  2. 在settings.py中启用中间件:
DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
PLAYWRIGHT_LAUNCH_OPTIONS = {
    "headless": True,
    "args": ["--disable-blink-features=AutomationControlled"],
}
PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT = 30000
  1. 编写爬虫:
import scrapy
from scrapy_playwright.page import PageCoroutine

class GoglobalSpider(scrapy.Spider):
    name = "goglobal"
    start_urls = ["https://goglobal.com/"]

    def start_requests(self):
        for url in self.start_urls:
            yield scrapy.Request(
                url,
                meta={
                    "playwright": True,
                    "playwright_page_coroutines": [
                        PageCoroutine("wait_for_selector", ".counter-number"),
                        PageCoroutine("evaluate", "window.scrollTo(0, document.body.scrollHeight)"),
                    ],
                },
            )

    def parse(self, response):
        counters = response.css(".counter-number::text").getall()
        yield {"counters": counters}

内容的提问来源于stack exchange,提问作者Adi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 15:40:20