You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup解析页面地址并解决URL拼接导致的连接错误

问题修复方案

错误根因

ConnectionError: HTTPConnectionPool(host='www.zrg74.ruhttp', port=80): Max retries exceeded with url: //zrg74.ru/sport/item/26982-xxx.html (Caused by NewConnectionError(...[Errno 11001] getaddrinfo failed))

报错的核心原因是URL拼接逻辑错误:你提取到的a标签href属性值为协议相对路径(格式如//zrg74.ru/xxx/xxx.html),或已经是完整的带http前缀的绝对URL,直接拼接固定前缀http://www.zrg74.ru会导致URL格式错乱,域名解析失败。

修复步骤

  • 引入Python标准库的urljoin工具自动处理URL拼接,无需手动拼接字符串,自动适配相对路径、协议相对路径、绝对路径等多种场景
  • 增加标签合法性校验,避免空标签引发的属性报错
  • 可按需统一域名格式,匹配你期望的输出样式(无www前缀)

修正后的代码片段

首先导入依赖:

from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
import re
from os import makedirs

核心逻辑修正如下:

# 建议指定解析器,避免环境差异导致的解析警告
soup = BeautifulSoup(page_text, 'html.parser')
posts_list = soup.find_all('div', {'class': 'jeg_post_excerpt'})
# 初始化集合存储结果,匹配你期望的输出格式
result_urls = set()

for p in posts_list:
    a_tag = p.find('a')
    # 提前校验a标签是否存在、是否有href属性,避免异常
    if not a_tag or 'href' not in a_tag.attrs:
        continue
    lnk = a_tag.attrs['href']
    title = re.sub('[^А-ЯЁа-яё0-9\s]', ' ', p.text)
    title = re.sub('\s\s+', ' ', title)
    
    # 用urljoin自动拼接,替代手动字符串拼接
    page_url = urljoin('http://www.zrg74.ru', lnk)
    # 统一域名格式,去掉www前缀匹配你的输出示例
    page_url = page_url.replace('http://www.zrg74.ru', 'http://zrg74.ru')
    
    # 简单校验URL合法性,过滤无效请求
    if not page_url.startswith('http'):
        continue
    # 加入结果集合
    result_urls.add(page_url)

    clean_path = '/'.join([d for d in page_url.split('/')[2:] if len(d) > 0])
    page_text = get_page_text(page_url, USER_AGENT)
    if page_text is None:
        continue
    dir_path = 'data/raw_pages/' + '/'.join(clean_path.split('/')[:-1])
    makedirs(dir_path, exist_ok=True)
    with open(dir_path + '/' + clean_path.split('/')[-1] + '.html', 'w', encoding='utf-8') as f:
        f.write(page_text)

# 输出你要求的格式
print(result_urls)

额外优化建议

  • get_page_text函数建议增加异常捕获逻辑,避免请求报错直接终止程序
  • 批量请求前可增加1-2秒的延时,避免触发站点反爬策略
  • 可先打印提取到的href和拼接后的page_url,确认格式正确后再执行批量采集

内容的提问来源于stack exchange,提问作者Alex_Kazantsev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 02:45:03