如何使用BeautifulSoup解析页面地址并解决URL拼接导致的连接错误
问题修复方案
错误根因
ConnectionError: HTTPConnectionPool(host='www.zrg74.ruhttp', port=80): Max retries exceeded with url: //zrg74.ru/sport/item/26982-xxx.html (Caused by NewConnectionError(...[Errno 11001] getaddrinfo failed))
报错的核心原因是URL拼接逻辑错误:你提取到的a标签href属性值为协议相对路径(格式如//zrg74.ru/xxx/xxx.html),或已经是完整的带http前缀的绝对URL,直接拼接固定前缀http://www.zrg74.ru会导致URL格式错乱,域名解析失败。
修复步骤
- 引入Python标准库的
urljoin工具自动处理URL拼接,无需手动拼接字符串,自动适配相对路径、协议相对路径、绝对路径等多种场景 - 增加标签合法性校验,避免空标签引发的属性报错
- 可按需统一域名格式,匹配你期望的输出样式(无www前缀)
修正后的代码片段
首先导入依赖:
from urllib.parse import urljoin import requests from bs4 import BeautifulSoup import re from os import makedirs
核心逻辑修正如下:
# 建议指定解析器,避免环境差异导致的解析警告 soup = BeautifulSoup(page_text, 'html.parser') posts_list = soup.find_all('div', {'class': 'jeg_post_excerpt'}) # 初始化集合存储结果,匹配你期望的输出格式 result_urls = set() for p in posts_list: a_tag = p.find('a') # 提前校验a标签是否存在、是否有href属性,避免异常 if not a_tag or 'href' not in a_tag.attrs: continue lnk = a_tag.attrs['href'] title = re.sub('[^А-ЯЁа-яё0-9\s]', ' ', p.text) title = re.sub('\s\s+', ' ', title) # 用urljoin自动拼接,替代手动字符串拼接 page_url = urljoin('http://www.zrg74.ru', lnk) # 统一域名格式,去掉www前缀匹配你的输出示例 page_url = page_url.replace('http://www.zrg74.ru', 'http://zrg74.ru') # 简单校验URL合法性,过滤无效请求 if not page_url.startswith('http'): continue # 加入结果集合 result_urls.add(page_url) clean_path = '/'.join([d for d in page_url.split('/')[2:] if len(d) > 0]) page_text = get_page_text(page_url, USER_AGENT) if page_text is None: continue dir_path = 'data/raw_pages/' + '/'.join(clean_path.split('/')[:-1]) makedirs(dir_path, exist_ok=True) with open(dir_path + '/' + clean_path.split('/')[-1] + '.html', 'w', encoding='utf-8') as f: f.write(page_text) # 输出你要求的格式 print(result_urls)
额外优化建议
get_page_text函数建议增加异常捕获逻辑,避免请求报错直接终止程序- 批量请求前可增加1-2秒的延时,避免触发站点反爬策略
- 可先打印提取到的
href和拼接后的page_url,确认格式正确后再执行批量采集
内容的提问来源于stack exchange,提问作者Alex_Kazantsev
相关产品推荐
相关产品推荐

