Python 3.9在Ubuntu中如何打开UTF-8编码文件名的文件?
解决Ubuntu下Python读取UTF-8编码非ASCII文件名的问题
问题背景
Python 3.9 + Ubuntu Linux环境,处理网站上传的UTF-8格式非ASCII文件名:
- 写入文件时,通过
naming.encode('utf-8', 'surrogateescape')生成路径(如directory/path/\xd8\xb9\xd8\xb1\xd8\xa8\xd9.txt),写入正常。 - 遍历目录读取时,调用
file.as_posix()打开文件报错'ascii' codec can't encode characters in position...;尝试转UTF-8编码也失败,且str(filepath)或filepath.as_posix()返回directory/path/????????.txt,修改系统locale为C.UTF-8后问题仍存在。
解决方案
1. 统一用pathlib处理路径(推荐)
Python 3的pathlib模块原生支持非ASCII路径,无需手动编码,避免字符串拼接的编码问题:
写入文件
from pathlib import Path # 用Path对象拼接路径,自动处理编码 file_path = Path(path) / filename with file_path.open('wb') as f: f.write(binary_data)
读取文件
from pathlib import Path target_dir = Path(path) # 遍历目录下的文件 for file_path in target_dir.iterdir(): # 直接用Path对象打开,自动处理路径编码 with file_path.open('rb') as f: data = f.read() # 后续数据处理逻辑
2. 直接使用字节路径打开
如果必须手动处理编码,读取时直接获取路径的字节形式,跳过字符串转码步骤:
from pathlib import Path target_dir = Path(path) for file_path in target_dir.iterdir(): # 用surrogateescape编码获取字节路径 bytes_path = file_path.as_posix().encode('utf-8', 'surrogateescape') with open(bytes_path, 'rb') as f: data = f.read() # 后续数据处理逻辑
3. 确认Python运行时编码
检查Python的文件系统编码是否为UTF-8:
import sys print(sys.getfilesystemencoding())
若输出不是utf-8,在运行脚本前设置环境变量:
export LC_ALL=C.UTF-8 export LANG=C.UTF-8
4. 上传环节的文件名处理
确保网站接收上传文件名时,已正确解码为UTF-8字符串,避免保留原始字节串导致路径异常(如Flask/Django框架中,上传文件的filename属性应为已解码的字符串)。
内容的提问来源于stack exchange,提问作者spospider
相关产品推荐
相关产品推荐

