网页抓取报错:Max retries exceeded with url 问题求助
公司环境下网页抓取失败,怀疑代理问题?
问题情况
跟随Keith Galli的网页抓取教程练习,目标网站为https://keithgalli.github.io/web-scraping/example.html,尝试抓取页面标题。教程中代码可成功运行,但在公司环境下执行失败;替换为公司内网网址则能正常运行,怀疑是代理问题,但不确定。
执行代码
# Load the webpage content r = requests.get("https://keithgalli.github.io/web-scraping/example.html") # Convert to a beautiful soup object soup = bs(r.content) # Print out our html print(soup.prettify())
报错信息
ConnectionError: HTTPSConnectionPool(host='keithgalli.github.io', port=443): Max retries exceeded with url: /web-scraping/example.html (Caused by NewConnectionError('<urllib3.connection.HTTPSConnection object at 0x0000026C42D80190>: Failed to establish a new connection: [Errno 11001] getaddrinfo failed'))
解决方案
从报错信息getaddrinfo failed来看,大概率是网络连接或代理配置问题,以下是几种处理方式:
1. 手动配置代理
如果知道公司代理地址和端口,直接在requests.get中传入proxies参数:
import requests from bs4 import BeautifulSoup as bs proxies = { "http": "http://你的代理地址:端口", "https": "https://你的代理地址:端口" } # 带代理请求目标网站 r = requests.get("https://keithgalli.github.io/web-scraping/example.html", proxies=proxies) soup = bs(r.content) print(soup.prettify())
2. 自动读取系统代理
让requests自动获取系统的代理配置,无需手动填写:
import requests from bs4 import BeautifulSoup as bs from requests.utils import get_environ_proxies url = "https://keithgalli.github.io/web-scraping/example.html" # 获取当前URL对应的系统代理 proxies = get_environ_proxies(url) r = requests.get(url, proxies=proxies) soup = bs(r.content) print(soup.prettify())
3. 排查DNS解析问题
getaddrinfo failed也可能是DNS无法解析目标域名,可以先在终端用ping keithgalli.github.io测试是否能正常解析。如果无法解析,可尝试手动指定DNS服务器,或者联系公司IT确认网络配置。
内容的提问来源于stack exchange,提问作者Yount Shi
相关产品推荐
相关产品推荐

