You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python urlopen()处理含波斯字符URL的编码解码问题求助

解决非ASCII URL适配urlopen()的编码问题

我来帮你搞定这个问题!从Chrome复制的带非ASCII字符的URL直接传给urlopen()会报错,因为urllib要求URL必须是符合标准的纯ASCII编码格式。下面是具体的解决思路和代码示例:

问题原因

Chrome浏览器会自动解码URL中的非ASCII字符显示给你,但urlopen()只接受RFC标准的ASCII编码URL,直接传入带非ASCII字符或未编码特殊字符(比如空格)的URL会抛出请求错误。

解决方案步骤

我们可以通过拆分URL、编码路径部分的方式,把非ASCII URL转换成合法格式,具体步骤如下:

  • 拆分URL结构:用urlparse把URL拆分成协议、域名、路径等部分,只对路径里的非ASCII/特殊字符编码
  • 编码路径部分:用quote()方法编码路径,保留必要的分隔符(比如/)
  • 重组URL:把编码后的路径和其他部分重新组合成合法URL
  • 分页URL处理:每次爬取下一页链接时,重复上述编码步骤,确保新URL也符合要求

完整代码示例

from urllib.request import urlopen, Request
from urllib.parse import quote, urlparse, urlunparse
from bs4 import BeautifulSoup

# 从Chrome复制的原始URL
original_url = 'https://www.sheypoor.com/%DA%A9%D9%85%D8%AF %D9%86%D9%88%D8%AC%D9%88%D8%A7%D9%86-34926671.html'

# 处理初始URL
parsed_url = urlparse(original_url)
# 编码路径部分,safe参数保留/不被编码
encoded_path = quote(parsed_url.path, safe='/')
# 重组合法的ASCII URL
current_url = urlunparse((
    parsed_url.scheme,
    parsed_url.netloc,
    encoded_path,
    parsed_url.params,
    parsed_url.query,
    parsed_url.fragment
))

# 循环爬取1-9页
for i in range(1, 10):
    # 添加User-Agent请求头,避免被网站反爬拦截
    req = Request(current_url, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'})
    try:
        html = urlopen(req)
    except Exception as e:
        print(f"第{i}页请求失败: {e}")
        break
    
    page = BeautifulSoup(html.read(), 'html.parser')
    
    # 这里需要根据页面实际结构调整选择器,找到下一页的链接
    next_page_tag = page.find('a', text='下一页')  # 示例选择器,需替换成实际页面的元素
    if not next_page_tag:
        print("已无下一页,停止爬取")
        break
    
    next_page_href = next_page_tag.get('href')
    # 处理下一页URL
    parsed_next = urlparse(next_page_href)
    # 如果下一页是相对路径,补全域名和协议
    next_scheme = parsed_next.scheme if parsed_next.scheme else parsed_url.scheme
    next_netloc = parsed_next.netloc if parsed_next.netloc else parsed_url.netloc
    encoded_next_path = quote(parsed_next.path, safe='/')
    
    current_url = urlunparse((
        next_scheme,
        next_netloc,
        encoded_next_path,
        parsed_next.params,
        parsed_next.query,
        parsed_next.fragment
    ))
    print(f"已准备爬取第{i+1}页: {current_url}")

关键说明

  • urlparse+urlunparse:拆分URL后只编码路径部分,不会破坏协议、域名等ASCII部分的结构
  • quote()的safe参数:设置为'/'确保路径分隔符不被编码,否则URL会失效
  • 添加User-Agent:很多网站会拦截无请求头的爬虫请求,模拟浏览器请求头能避免大部分拦截
  • 相对路径处理:如果下一页链接是相对路径,要补全协议和域名,否则urlopen无法识别

内容的提问来源于stack exchange,提问作者Homayoun Soleimani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:07:16