爬取网站URL遇UnicodeError求助:label空或过长
问题描述
爬取http://www.nuigalway.ie时触发以下错误:
Getting this error - hostIP = socket.gethostbyname(hostonly) UnicodeError: encoding with 'idna' codec failed (UnicodeError: label empty or too long)
实现代码如下:
from bs4 import BeautifulSoup import requests from urllib.parse import urljoin, quote def scrape(site): visited = set() queue = [site] while queue: current_url = queue.pop(0) if current_url in visited: continue visited.add(current_url) try: r = requests.get(current_url, timeout=5) except requests.exceptions.RequestException as e: print(f"Error connecting to {current_url}: {e}") continue soup = BeautifulSoup(r.text, "html.parser") for link in soup.find_all("a"): href = link.get("href") if href is not None: full_url = urljoin(site, quote(href, safe='/:?=&')) if site in full_url and full_url not in visited: queue.append(full_url) print(full_url) if __name__ == "__main__": site = "http://www.nuigalway.ie" scrape(site)
解决方法
这个错误是因为生成的URL包含不符合IDNA编码规范的域名片段(比如空标签、过长的域名部分),导致socket解析域名失败。可以通过以下步骤修复:
1. 过滤无效URL片段
在处理href时,先跳过明显无效的链接(比如空字符串、仅含特殊字符的链接),避免生成非法URL。
2. 修正URL处理逻辑
不要先对href做quote再拼接,应该先用urljoin生成完整URL,再对URL的查询参数等部分做合理编码,避免重复编码破坏域名结构。同时用urlparse验证URL的合法性。
3. 添加异常捕获
在处理和请求URL时,捕获UnicodeError和其他URL相关异常,跳过有问题的链接,保证爬虫持续运行。
修改后的代码
from bs4 import BeautifulSoup import requests from urllib.parse import urljoin, urlparse, quote def is_valid_url(url, base_domain): # 验证URL是否合法且属于目标域名 parsed = urlparse(url) # 检查是否有合法的协议和域名,且域名包含目标域名 if not parsed.scheme or not parsed.netloc: return False return base_domain in parsed.netloc def scrape(site): visited = set() queue = [site] base_domain = urlparse(site).netloc # 提取目标域名 while queue: current_url = queue.pop(0) if current_url in visited: continue visited.add(current_url) try: r = requests.get(current_url, timeout=5) except (requests.exceptions.RequestException, UnicodeError) as e: print(f"跳过无效URL {current_url}: {e}") continue soup = BeautifulSoup(r.text, "html.parser") for link in soup.find_all("a"): href = link.get("href") if not href or href.strip() == "": continue # 跳过空链接 # 先拼接完整URL,再处理查询参数编码 full_url = urljoin(site, href) parsed_url = urlparse(full_url) # 仅编码查询参数部分,避免破坏域名 encoded_query = quote(parsed_url.query, safe='=&') full_url = parsed_url._replace(query=encoded_query).geturl() if is_valid_url(full_url, base_domain) and full_url not in visited: try: # 提前验证域名是否可被IDNA编码 parsed_url.netloc.encode('idna') queue.append(full_url) print(full_url) except UnicodeError as e: print(f"跳过域名非法的URL {full_url}: {e}") if __name__ == "__main__": site = "http://www.nuigalway.ie" scrape(site)
关键修改说明
- 新增
is_valid_url函数,过滤掉无协议、无域名的非法链接,确保只爬取目标域名下的内容 - 调整URL编码顺序:先拼接再编码查询参数,避免破坏域名结构
- 新增域名IDNA编码预验证,提前过滤无法解析的异常域名
- 扩展异常捕获范围,跳过所有异常URL,保证爬虫稳定运行
内容的提问来源于stack exchange,提问作者Abhidith N Shetty
相关产品推荐
相关产品推荐

