You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取报错:Max retries exceeded with url 问题求助

公司环境下网页抓取失败,怀疑代理问题?

问题情况

跟随Keith Galli的网页抓取教程练习,目标网站为https://keithgalli.github.io/web-scraping/example.html,尝试抓取页面标题。教程中代码可成功运行,但在公司环境下执行失败;替换为公司内网网址则能正常运行,怀疑是代理问题,但不确定。

执行代码

# Load the webpage content
r = requests.get("https://keithgalli.github.io/web-scraping/example.html")

# Convert to a beautiful soup object
soup = bs(r.content)

# Print out our html
print(soup.prettify())

报错信息

ConnectionError: HTTPSConnectionPool(host='keithgalli.github.io', port=443): Max retries exceeded with url: /web-scraping/example.html (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x0000026C42D80190>: Failed to establish a new connection: [Errno 11001] getaddrinfo failed'))

解决方案

从报错信息getaddrinfo failed来看,大概率是网络连接或代理配置问题,以下是几种处理方式:

1. 手动配置代理

如果知道公司代理地址和端口,直接在requests.get中传入proxies参数:

import requests
from bs4 import BeautifulSoup as bs

proxies = {
    "http": "http://你的代理地址:端口",
    "https": "https://你的代理地址:端口"
}

# 带代理请求目标网站
r = requests.get("https://keithgalli.github.io/web-scraping/example.html", proxies=proxies)
soup = bs(r.content)
print(soup.prettify())

2. 自动读取系统代理

让requests自动获取系统的代理配置,无需手动填写:

import requests
from bs4 import BeautifulSoup as bs
from requests.utils import get_environ_proxies

url = "https://keithgalli.github.io/web-scraping/example.html"
# 获取当前URL对应的系统代理
proxies = get_environ_proxies(url)
r = requests.get(url, proxies=proxies)
soup = bs(r.content)
print(soup.prettify())

3. 排查DNS解析问题

getaddrinfo failed也可能是DNS无法解析目标域名,可以先在终端用ping keithgalli.github.io测试是否能正常解析。如果无法解析,可尝试手动指定DNS服务器,或者联系公司IT确认网络配置。

内容的提问来源于stack exchange,提问作者Yount Shi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 15:25:14