使用requests_html提取JS内容失败的问题及解决方案咨询
爬虫动态内容提取问题解答
问题背景
开发网络爬虫时,用Beautiful Soup处理静态HTML正常,用requests_html提取动态JS内容在多数网站有效,但爬取https://goglobal.com/时,无法获取“100+ countries”“2500+ employees”“3 Billion dollars saved...”区域的数值,测试脚本已尝试延长等待时间但无效。
测试脚本:
from requests_html import HTMLSession import time session = HTMLSession() url = "https://goglobal.com/" r = session.get(url) r.html.render(wait=10) time.sleep(10) print(r.html.html)
1. 为何内容无法正确加载?
- 反爬检测:网站可能识别出
requests_html依赖的无头Pyppeteer浏览器特征,限制了动态内容渲染,这类反爬会针对navigator.webdriver等属性做拦截。 - 触发条件未满足:这些统计数字需要页面滚动到对应区域才会触发JS渲染逻辑,
requests_html默认只渲染首屏内容,无滚动动作则不会加载目标内容。 - JS环境缺失:
requests_html的Pyppeteer环境可能缺少网站所需的部分JS API或特性,导致统计数字的渲染代码执行失败。
2. 能否继续使用requests_html解决此问题?
可以尝试以下调整方案,仍有概率解决问题:
- 添加真实User-Agent:模拟普通浏览器请求,避免被识别为爬虫
- 设置滚动触发:通过
scrolldown参数让页面滚动,触发内容加载 - 调整等待逻辑:延长渲染等待时间,确保JS完全执行
调整后的测试脚本:
from requests_html import HTMLSession session = HTMLSession() url = "https://goglobal.com/" # 自定义浏览器User-Agent headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } r = session.get(url, headers=headers) # 滚动2次触发内容加载,每次间隔2秒 r.html.render(wait=5, scrolldown=2, sleep=2) # 提取目标元素 counters = r.html.find('.counter-number') for counter in counters: print(counter.text)
如果上述调整仍无效,说明网站反爬机制针对Pyppeteer做了强拦截,requests_html的局限性会凸显,难以绕过。
3. 是否可通过selenium、playwright或scrapy解决此问题?
这三个工具都可以解决,具体方案如下:
- Playwright:对无头浏览器的伪装性最好,内置反爬规避特性,代码实现简洁:
from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page(user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') page.goto("https://goglobal.com/", wait_until="networkidle") # 滚动到统计区域触发渲染 page.locator('.counter-number').first.scroll_into_view_if_needed() page.wait_for_selector('.counter-number', state='visible') counters = page.locator('.counter-number').all_text_contents() print(counters) browser.close()
- Selenium:兼容性强,可通过修改浏览器参数规避自动化检测:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = Options() options.add_argument('--headless=new') options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # 禁用自动化检测特征 options.add_argument('--disable-blink-features=AutomationControlled') driver = webdriver.Chrome(options=options) driver.get("https://goglobal.com/") # 等待元素加载并滚动到目标区域 wait = WebDriverWait(driver, 15) counters = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'counter-number'))) for counter in counters: driver.execute_script("arguments[0].scrollIntoView();", counter) print(counter.text) driver.quit()
- Scrapy:本身是静态爬虫框架,但可通过
scrapy-playwright中间件集成Playwright处理动态内容,适合大规模爬取场景:
- 安装依赖:
pip install scrapy-playwright - 在
settings.py中启用中间件:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, "args": ["--disable-blink-features=AutomationControlled"], } PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT = 30000
- 编写爬虫:
import scrapy from scrapy_playwright.page import PageCoroutine class GoglobalSpider(scrapy.Spider): name = "goglobal" start_urls = ["https://goglobal.com/"] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={ "playwright": True, "playwright_page_coroutines": [ PageCoroutine("wait_for_selector", ".counter-number"), PageCoroutine("evaluate", "window.scrollTo(0, document.body.scrollHeight)"), ], }, ) def parse(self, response): counters = response.css(".counter-number::text").getall() yield {"counters": counters}
内容的提问来源于stack exchange,提问作者Adi
相关产品推荐
相关产品推荐

