Python urllib下载含特殊字符URL/文件时遇UnicodeEncodeError问题
解决含挪威语特殊字符的Azure Blob URL批量下载UnicodeEncodeError问题
问题
批量下载Excel中指向Azure Blob的视频URL时,URL或文件名含æ、ø、å等挪威语特殊字符,触发编码错误:
UnicodeEncodeError: 'ascii' codec can't encode character '\xf8' in position 66: ordinal not in range(128)
原代码
import pandas as pd import os import urllib.request import urllib.error df = pd.read_excel('MC-Redo.xlsx') df_column = df.iloc[:,1] def getVideo(): for value in df_column: if "https://pinnacle.blob.core.windows.net/" in str(value) and " " not in str(value): if value.find('/'): fileName = value.rsplit('/', 1)[1] if not os.path.exists(path+fileName): urllib.request.urlretrieve(value, fileName) if "https://pinnacle.blob.core.windows.net/" in str(value) and " " in str(value): newUrl = value.replace(' ', '%20') if newUrl.find('/'): fileName = newUrl.rsplit('/', 1)[1] if not os.path.exists(path+fileName): try: urllib.request.urlretrieve(newUrl, fileName) except urllib.error.HTTPError as e: if e.code != 200: continue
解决方案
核心修改点
- URL全字符编码:用
urllib.parse.quote()对URL路径部分进行编码,覆盖所有非ASCII字符(包括æ、ø、å和空格),而非仅替换空格。 - 文件名解码还原:用
urllib.parse.unquote()将编码后的文件名解码回原始字符,保证本地文件名显示正确。 - 路径安全拼接:使用
os.path.join()处理保存路径,避免手动拼接字符串导致的编码冲突。 - 修复逻辑漏洞:替换
value.find('/')为'/' in value,避免因索引为0导致的逻辑判断错误。
修改后代码
import pandas as pd import os import urllib.request import urllib.error from urllib.parse import quote, unquote df = pd.read_excel('MC-Redo.xlsx') df_column = df.iloc[:,1] path = "./downloads/" # 替换为你的目标保存路径 def getVideo(): # 自动创建保存目录(不存在时) os.makedirs(path, exist_ok=True) for value in df_column: url_str = str(value).strip() # 跳过非目标URL if "https://pinnacle.blob.core.windows.net/" not in url_str: continue # 拆分URL并编码路径部分,保留/作为分隔符 protocol, rest = url_str.split('://', 1) domain, path_part = rest.split('/', 1) encoded_path = quote(path_part, safe='/') encoded_url = f"{protocol}://{domain}/{encoded_path}" # 提取并解码文件名 file_name = unquote(encoded_path.rsplit('/', 1)[1]) save_path = os.path.join(path, file_name) if not os.path.exists(save_path): try: urllib.request.urlretrieve(encoded_url, save_path) except urllib.error.HTTPError as e: if e.code != 200: continue getVideo()
说明
- 编码URL时,
safe='/'确保路径分隔符不被编码,维持URL结构正确性。 os.makedirs(path, exist_ok=True)避免因保存目录不存在引发的报错。- 统一处理所有含特殊字符的URL,简化代码逻辑。
内容的提问来源于stack exchange,提问作者Alexander Brandhaug
相关产品推荐
相关产品推荐

