Python网页爬取代码报错:域名解析失败问题求助
解决BBC新闻爬虫的连接错误问题
问题代码
import pyperclip import requests from bs4 import BeautifulSoup base_url = "https://www.bbc.com" url = base_url + "/news/world" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') articles = soup.find_all('div', class_='gs-c-promo-body') text = '' for article in articles: headline = article.find('h3', class_='gs-c-promo-heading__title') if headline: text += headline.text + '\n' summary = article.find('p', class_='gs-c-promo-summary') if summary: text += summary.text + '\n' link = article.find('a', class_='gs-c-promo-heading') if link: href = link['href'] if href.startswith('//'): article_url = 'https:' + href else: article_url = base_url + href article_response = requests.get(article_url) article_soup = BeautifulSoup(article_response.text, 'html.parser') article_text = article_soup.find('div', class_='story-body__inner') if article_text: text += article_text.get_text() + '\n\n' pyperclip.copy(text)
报错信息
Traceback (most recent call last): File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 200, in _new_conn sock = connection.create_connection( File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\util\connection.py", line 60, in create_connection for res in socket.getaddrinfo(host, port, family, socket.SOCK_STREAM): File "C:\Users\msala\AppData\Local\Programs\Python\Python39\lib\socket.py", line 954, in getaddrinfo for res in _socket.getaddrinfo(host, port, family, type, proto, flags): socket.gaierror: [Errno 11001] getaddrinfo failed The above exception was the direct cause of the following exception: Traceback (most recent call last): File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 790, in urlopen response = self._make_request( File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 491, in _make_request raise new_e File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 467, in _make_request self._validate_conn(conn) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 1092, in _validate_conn conn.connect() File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 604, in connect self.sock = sock = self._new_conn() File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 207, in _new_conn raise NameResolutionError(self.host, self, e) from e urllib3.exceptions.NameResolutionError: <urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed) The above exception was the direct cause of the following exception: Traceback (most recent call last): File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\adapters.py", line 486, in send resp = conn.urlopen( File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 844, in urlopen retries = retries.increment( File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\util\retry.py", line 515, in increment raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] urllib3.exceptions.MaxRetryError: HTTPSConnectionPool(host='www.bbc.comhttps', port=443): Max retries exceeded with url: //www.bbc.com/future/article/20230512-eurovision-why-some-countries-vote-for-each-other (Caused by NameResolutionError("<urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed)")) During handling of the above exception, another exception occurred: Traceback (most recent call last): File "C:\Users\msala\PycharmProjects\pythonProject1\main.py", line 26, in <module> article_response = requests.get(article_url) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\api.py", line 73, in get return request("get", url, params=params, **kwargs) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\api.py", line 59, in request return session.request(method=method, url=url, **kwargs) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\sessions.py", line 587, in request resp = self.send(prep, **send_kwargs) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\sessions.py", line 701, in send r = adapter.send(request, **kwargs) File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\adapters.py", line 519, in send raise ConnectionError(e, request=request) requests.exceptions.ConnectionError: HTTPSConnectionPool(host='www.bbc.comhttps', port=443): Max retries exceeded with url: //www.bbc.com/future/article/20230512-eurovision-why-some-countries-vote-for-each-other (Caused by NameResolutionError("<urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed)")) Process finished with exit code 1
错误原因
从报错信息里的host='www.bbc.comhttps'可以看出,URL拼接逻辑存在问题:部分文章的href已经是完整的HTTPS链接(比如https://www.bbc.com/future/xxx),但代码直接将base_url与该完整链接拼接,生成了https://www.bbc.comhttps://www.bbc.com/future/xxx这类错误URL,进而触发域名解析失败。
修复方案
- 使用
urllib.parse.urljoin自动处理URL拼接,它能根据基础URL和相对/绝对URL生成正确的完整链接。 - 添加请求头模拟浏览器,避免被BBC反爬机制拦截。
- 增加异常处理,确保单个文章请求失败不会导致整个程序终止。
修改后的完整代码
import pyperclip import requests from bs4 import BeautifulSoup from urllib.parse import urljoin base_url = "https://www.bbc.com" url = urljoin(base_url, "/news/world") # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: response = requests.get(url, headers=headers) response.raise_for_status() # 检查请求是否成功 soup = BeautifulSoup(response.text, 'html.parser') articles = soup.find_all('div', class_='gs-c-promo-body') text = '' for article in articles: # 提取标题 headline = article.find('h3', class_='gs-c-promo-heading__title') if headline: text += headline.text.strip() + '\n' # 提取摘要 summary = article.find('p', class_='gs-c-promo-summary') if summary: text += summary.text.strip() + '\n' # 处理文章链接 link = article.find('a', class_='gs-c-promo-heading') if link and 'href' in link.attrs: href = link['href'] article_url = urljoin(base_url, href) # 自动处理相对/绝对URL try: article_response = requests.get(article_url, headers=headers) article_response.raise_for_status() article_soup = BeautifulSoup(article_response.text, 'html.parser') # BBC页面结构可能更新,兼容新旧正文选择器 article_text = article_soup.find('div', class_='ssrcss-11r1m41-RichTextContainer') or article_soup.find('div', class_='story-body__inner') if article_text: text += article_text.get_text(separator='\n').strip() + '\n\n' except requests.exceptions.RequestException as e: text += f"获取文章内容失败: {str(e)}\n\n" continue pyperclip.copy(text) print("内容已复制到剪贴板") except requests.exceptions.RequestException as e: print(f"请求主页面失败: {str(e)}")
额外说明
- BBC页面结构可能随时更新,若无法提取正文,需检查页面元素的class是否变化,调整对应选择器。
- 频繁爬取可能触发反爬限制,建议添加适当延迟(如
time.sleep(1))。
内容的提问来源于stack exchange,提问作者veryNoobProgrammer
相关产品推荐
相关产品推荐

