You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取网站URL遇UnicodeError求助:label空或过长

问题描述

爬取http://www.nuigalway.ie时触发以下错误:

Getting this error - 
    hostIP = socket.gethostbyname(hostonly)
UnicodeError: encoding with 'idna' codec failed (UnicodeError: label empty or too long)

实现代码如下:

from bs4 import BeautifulSoup
import requests
from urllib.parse import urljoin, quote

def scrape(site):
    visited = set()
    queue = [site]

    while queue:
        current_url = queue.pop(0)
        if current_url in visited:
            continue

        visited.add(current_url)

        try:
            r = requests.get(current_url, timeout=5)
        except requests.exceptions.RequestException as e:
            print(f"Error connecting to {current_url}: {e}")
            continue

        soup = BeautifulSoup(r.text, "html.parser")
        for link in soup.find_all("a"):
            href = link.get("href")
            if href is not None:
                full_url = urljoin(site, quote(href, safe='/:?=&'))
                if site in full_url and full_url not in visited:
                    queue.append(full_url)
                    print(full_url)

if __name__ == "__main__":
    site = "http://www.nuigalway.ie"
    scrape(site)
解决方法

这个错误是因为生成的URL包含不符合IDNA编码规范的域名片段(比如空标签、过长的域名部分),导致socket解析域名失败。可以通过以下步骤修复:

1. 过滤无效URL片段

在处理href时,先跳过明显无效的链接(比如空字符串、仅含特殊字符的链接),避免生成非法URL。

2. 修正URL处理逻辑

不要先对href做quote再拼接,应该先用urljoin生成完整URL,再对URL的查询参数等部分做合理编码,避免重复编码破坏域名结构。同时用urlparse验证URL的合法性。

3. 添加异常捕获

在处理和请求URL时,捕获UnicodeError和其他URL相关异常,跳过有问题的链接,保证爬虫持续运行。

修改后的代码

from bs4 import BeautifulSoup
import requests
from urllib.parse import urljoin, urlparse, quote

def is_valid_url(url, base_domain):
    # 验证URL是否合法且属于目标域名
    parsed = urlparse(url)
    # 检查是否有合法的协议和域名,且域名包含目标域名
    if not parsed.scheme or not parsed.netloc:
        return False
    return base_domain in parsed.netloc

def scrape(site):
    visited = set()
    queue = [site]
    base_domain = urlparse(site).netloc  # 提取目标域名

    while queue:
        current_url = queue.pop(0)
        if current_url in visited:
            continue

        visited.add(current_url)

        try:
            r = requests.get(current_url, timeout=5)
        except (requests.exceptions.RequestException, UnicodeError) as e:
            print(f"跳过无效URL {current_url}: {e}")
            continue

        soup = BeautifulSoup(r.text, "html.parser")
        for link in soup.find_all("a"):
            href = link.get("href")
            if not href or href.strip() == "":
                continue  # 跳过空链接

            # 先拼接完整URL,再处理查询参数编码
            full_url = urljoin(site, href)
            parsed_url = urlparse(full_url)
            # 仅编码查询参数部分,避免破坏域名
            encoded_query = quote(parsed_url.query, safe='=&')
            full_url = parsed_url._replace(query=encoded_query).geturl()

            if is_valid_url(full_url, base_domain) and full_url not in visited:
                try:
                    # 提前验证域名是否可被IDNA编码
                    parsed_url.netloc.encode('idna')
                    queue.append(full_url)
                    print(full_url)
                except UnicodeError as e:
                    print(f"跳过域名非法的URL {full_url}: {e}")

if __name__ == "__main__":
    site = "http://www.nuigalway.ie"
    scrape(site)

关键修改说明

  • 新增is_valid_url函数,过滤掉无协议、无域名的非法链接,确保只爬取目标域名下的内容
  • 调整URL编码顺序:先拼接再编码查询参数,避免破坏域名结构
  • 新增域名IDNA编码预验证,提前过滤无法解析的异常域名
  • 扩展异常捕获范围,跳过所有异常URL,保证爬虫稳定运行

内容的提问来源于stack exchange,提问作者Abhidith N Shetty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 16:33:15