You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取代码报错:域名解析失败问题求助

解决BBC新闻爬虫的连接错误问题

问题代码

import pyperclip
import requests
from bs4 import BeautifulSoup

base_url = "https://www.bbc.com"
url = base_url + "/news/world"
response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')
articles = soup.find_all('div', class_='gs-c-promo-body')
text = ''
for article in articles:
    headline = article.find('h3', class_='gs-c-promo-heading__title')
    if headline:
        text += headline.text + '\n'
    summary = article.find('p', class_='gs-c-promo-summary')
    if summary:
        text += summary.text + '\n'
    link = article.find('a', class_='gs-c-promo-heading')
    if link:
        href = link['href']
        if href.startswith('//'):
            article_url = 'https:' + href
        else:
            article_url = base_url + href
        article_response = requests.get(article_url)
        article_soup = BeautifulSoup(article_response.text, 'html.parser')
        article_text = article_soup.find('div', class_='story-body__inner')
        if article_text:
            text += article_text.get_text() + '\n\n'
pyperclip.copy(text)

报错信息

Traceback (most recent call last):
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 200, in _new_conn
    sock = connection.create_connection(
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\util\connection.py", line 60, in create_connection
    for res in socket.getaddrinfo(host, port, family, socket.SOCK_STREAM):
  File "C:\Users\msala\AppData\Local\Programs\Python\Python39\lib\socket.py", line 954, in getaddrinfo
    for res in _socket.getaddrinfo(host, port, family, type, proto, flags):
socket.gaierror: [Errno 11001] getaddrinfo failed

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 790, in urlopen
    response = self._make_request(
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 491, in _make_request
    raise new_e
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 467, in _make_request
    self._validate_conn(conn)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 1092, in _validate_conn
    conn.connect()
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 604, in connect
    self.sock = sock = self._new_conn()
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connection.py", line 207, in _new_conn
    raise NameResolutionError(self.host, self, e) from e
urllib3.exceptions.NameResolutionError: <urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed)

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\adapters.py", line 486, in send
    resp = conn.urlopen(
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\connectionpool.py", line 844, in urlopen
    retries = retries.increment(
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\urllib3\util\retry.py", line 515, in increment
    raise MaxRetryError(_pool, url, reason) from reason  # type: ignore[arg-type]
urllib3.exceptions.MaxRetryError: HTTPSConnectionPool(host='www.bbc.comhttps', port=443): Max retries exceeded with url: //www.bbc.com/future/article/20230512-eurovision-why-some-countries-vote-for-each-other (Caused by NameResolutionError("<urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed)"))

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
  File "C:\Users\msala\PycharmProjects\pythonProject1\main.py", line 26, in <module>
    article_response = requests.get(article_url)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\api.py", line 73, in get
    return request("get", url, params=params, **kwargs)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\api.py", line 59, in request
    return session.request(method=method, url=url, **kwargs)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\sessions.py", line 587, in request
    resp = self.send(prep, **send_kwargs)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\sessions.py", line 701, in send
    r = adapter.send(request, **kwargs)
  File "C:\Users\msala\PycharmProjects\learnPython\venv\pythonProject1\lib\site-packages\requests\adapters.py", line 519, in send
    raise ConnectionError(e, request=request)
requests.exceptions.ConnectionError: HTTPSConnectionPool(host='www.bbc.comhttps', port=443): Max retries exceeded with url: //www.bbc.com/future/article/20230512-eurovision-why-some-countries-vote-for-each-other (Caused by NameResolutionError("<urllib3.connection.HTTPSConnection object at 0x000001AF8CB443D0>: Failed to resolve 'www.bbc.comhttps' ([Errno 11001] getaddrinfo failed)"))

Process finished with exit code 1

错误原因

从报错信息里的host='www.bbc.comhttps'可以看出,URL拼接逻辑存在问题:部分文章的href已经是完整的HTTPS链接(比如https://www.bbc.com/future/xxx),但代码直接将base_url与该完整链接拼接,生成了https://www.bbc.comhttps://www.bbc.com/future/xxx这类错误URL,进而触发域名解析失败。

修复方案

  1. 使用urllib.parse.urljoin自动处理URL拼接,它能根据基础URL和相对/绝对URL生成正确的完整链接。
  2. 添加请求头模拟浏览器,避免被BBC反爬机制拦截。
  3. 增加异常处理,确保单个文章请求失败不会导致整个程序终止。

修改后的完整代码

import pyperclip
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

base_url = "https://www.bbc.com"
url = urljoin(base_url, "/news/world")

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

try:
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # 检查请求是否成功
    soup = BeautifulSoup(response.text, 'html.parser')
    articles = soup.find_all('div', class_='gs-c-promo-body')
    text = ''
    
    for article in articles:
        # 提取标题
        headline = article.find('h3', class_='gs-c-promo-heading__title')
        if headline:
            text += headline.text.strip() + '\n'
        
        # 提取摘要
        summary = article.find('p', class_='gs-c-promo-summary')
        if summary:
            text += summary.text.strip() + '\n'
        
        # 处理文章链接
        link = article.find('a', class_='gs-c-promo-heading')
        if link and 'href' in link.attrs:
            href = link['href']
            article_url = urljoin(base_url, href)  # 自动处理相对/绝对URL
            
            try:
                article_response = requests.get(article_url, headers=headers)
                article_response.raise_for_status()
                article_soup = BeautifulSoup(article_response.text, 'html.parser')
                # BBC页面结构可能更新,兼容新旧正文选择器
                article_text = article_soup.find('div', class_='ssrcss-11r1m41-RichTextContainer') or article_soup.find('div', class_='story-body__inner')
                
                if article_text:
                    text += article_text.get_text(separator='\n').strip() + '\n\n'
            except requests.exceptions.RequestException as e:
                text += f"获取文章内容失败: {str(e)}\n\n"
                continue
    
    pyperclip.copy(text)
    print("内容已复制到剪贴板")

except requests.exceptions.RequestException as e:
    print(f"请求主页面失败: {str(e)}")

额外说明

  • BBC页面结构可能随时更新,若无法提取正文,需检查页面元素的class是否变化,调整对应选择器。
  • 频繁爬取可能触发反爬限制,建议添加适当延迟(如time.sleep(1))。

内容的提问来源于stack exchange,提问作者veryNoobProgrammer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:07:36