You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python urllib下载含特殊字符URL/文件时遇UnicodeEncodeError问题

解决含挪威语特殊字符的Azure Blob URL批量下载UnicodeEncodeError问题

问题

批量下载Excel中指向Azure Blob的视频URL时,URL或文件名含æ、ø、å等挪威语特殊字符,触发编码错误:

UnicodeEncodeError: 'ascii' codec can't encode character '\xf8' in position 66: ordinal not in range(128)

原代码

import pandas as pd
import os
import urllib.request
import urllib.error

df = pd.read_excel('MC-Redo.xlsx')
df_column = df.iloc[:,1]

def getVideo():
    for value in df_column:
        if "https://pinnacle.blob.core.windows.net/" in str(value) and " " not in str(value):
            
            if value.find('/'):
                fileName = value.rsplit('/', 1)[1]
                if not os.path.exists(path+fileName):
                    urllib.request.urlretrieve(value, fileName)

        if "https://pinnacle.blob.core.windows.net/" in str(value) and " " in str(value):
            newUrl = value.replace(' ', '%20')
            if newUrl.find('/'):
                fileName = newUrl.rsplit('/', 1)[1]
                if not os.path.exists(path+fileName):
                    try: 
                        urllib.request.urlretrieve(newUrl, fileName)
                    except urllib.error.HTTPError as e:
                        if e.code != 200:
                            continue

解决方案

核心修改点

  1. URL全字符编码:用urllib.parse.quote()对URL路径部分进行编码,覆盖所有非ASCII字符(包括æ、ø、å和空格),而非仅替换空格。
  2. 文件名解码还原:用urllib.parse.unquote()将编码后的文件名解码回原始字符,保证本地文件名显示正确。
  3. 路径安全拼接:使用os.path.join()处理保存路径,避免手动拼接字符串导致的编码冲突。
  4. 修复逻辑漏洞:替换value.find('/')为'/' in value,避免因索引为0导致的逻辑判断错误。

修改后代码

import pandas as pd
import os
import urllib.request
import urllib.error
from urllib.parse import quote, unquote

df = pd.read_excel('MC-Redo.xlsx')
df_column = df.iloc[:,1]
path = "./downloads/"  # 替换为你的目标保存路径

def getVideo():
    # 自动创建保存目录(不存在时)
    os.makedirs(path, exist_ok=True)
    
    for value in df_column:
        url_str = str(value).strip()
        # 跳过非目标URL
        if "https://pinnacle.blob.core.windows.net/" not in url_str:
            continue
            
        # 拆分URL并编码路径部分,保留/作为分隔符
        protocol, rest = url_str.split('://', 1)
        domain, path_part = rest.split('/', 1)
        encoded_path = quote(path_part, safe='/')
        encoded_url = f"{protocol}://{domain}/{encoded_path}"
        
        # 提取并解码文件名
        file_name = unquote(encoded_path.rsplit('/', 1)[1])
        save_path = os.path.join(path, file_name)
        
        if not os.path.exists(save_path):
            try:
                urllib.request.urlretrieve(encoded_url, save_path)
            except urllib.error.HTTPError as e:
                if e.code != 200:
                    continue

getVideo()

说明

  • 编码URL时,safe='/'确保路径分隔符不被编码,维持URL结构正确性。
  • os.makedirs(path, exist_ok=True)避免因保存目录不存在引发的报错。
  • 统一处理所有含特殊字符的URL,简化代码逻辑。

内容的提问来源于stack exchange,提问作者Alexander Brandhaug

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 16:10:33