You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests_html爬取网页时遭遇pyppeteer.TimeoutError求助

Fixing Navigation Timeout Exceeded Error with requests_html

Hey there, sorry you're hitting this timeout snag when trying to render the Beerwulf page. Let's break down what's going on and fix it step by step.

The error pyppeteer.errors.TimeoutError: Navigation Timeout Exceeded: 8000 ms exceeded happens because the pyppeteer tool (which requests_html uses under the hood to load JavaScript-heavy pages) is running out of time waiting for the page to fully load. Here are actionable solutions you can try:

1. Increase the Timeout Duration

The simplest fix is to give the page more breathing room by extending the timeout value in the render() method. The default is 8 seconds, so bump it up to something like 20 seconds:

from requests_html import HTMLSession
url = 'https://www.beerwulf.com/en-gb/c/beers?segment=Beers&catalogCode=Beer_1'
s = HTMLSession()
r = s.get(url)
# Add timeout parameter (value is in seconds)
r.html.render(sleep=1, timeout=20)
print(r.status_code)

2. Wait for a Specific Element to Load

Instead of waiting a fixed amount of time, tell requests_html to hold off until a critical element on the page appears. This is more efficient because it stops waiting as soon as the content you care about is ready. For example, if beer product cards use the class product-item, use the wait parameter:

r.html.render(sleep=1, timeout=20, wait='.product-item')

3. Disable Image Loading to Speed Up Page Load

Images are often the slowest part of a page to load. Disabling them can cut down load time drastically. You can either inject a small script to hide images or configure the browser to skip loading them entirely:

Option A: Inject a script to hide images

r.html.render(
    sleep=1,
    timeout=20,
    script="""
        // Hide all images to reduce load time
        document.querySelectorAll('img').forEach(img => img.style.display = 'none');
    """
)

Option B: Configure the browser to disable images

from requests_html import HTMLSession

# Create session with browser args to disable images
s = HTMLSession()
s.browser_args.extend([
    "--no-sandbox",
    "--disable-setuid-sandbox",
    "--disable-images"
])

url = 'https://www.beerwulf.com/en-gb/c/beers?segment=Beers&catalogCode=Beer_1'
r = s.get(url)
r.html.render(sleep=1, timeout=20)
print(r.status_code)

4. Mimic a Real Browser to Bypass Anti-Scraping Measures

Some sites block headless browsers (which pyppeteer uses by default). Adding a realistic user-agent and headers can help you fly under the radar:

from requests_html import HTMLSession

url = 'https://www.beerwulf.com/en-gb/c/beers?segment=Beers&catalogCode=Beer_1'
s = HTMLSession()

# Use a real user-agent header
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
}

r = s.get(url, headers=headers)
# You can also pass the user-agent directly to render()
r.html.render(sleep=1, timeout=20, user_agent=headers['User-Agent'])
print(r.status_code)

Start with the first solution (increasing timeout) and work your way down if needed. Most of the time, a longer timeout combined with waiting for a key element will resolve the issue.

内容的提问来源于stack exchange,提问作者Rehman Yz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:17:31