You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.9在Ubuntu中如何打开UTF-8编码文件名的文件?

解决Ubuntu下Python读取UTF-8编码非ASCII文件名的问题

问题背景

Python 3.9 + Ubuntu Linux环境,处理网站上传的UTF-8格式非ASCII文件名:

  • 写入文件时,通过naming.encode('utf-8', 'surrogateescape')生成路径(如directory/path/\xd8\xb9\xd8\xb1\xd8\xa8\xd9.txt),写入正常。
  • 遍历目录读取时,调用file.as_posix()打开文件报错'ascii' codec can't encode characters in position...;尝试转UTF-8编码也失败,且str(filepath)或filepath.as_posix()返回directory/path/????????.txt,修改系统locale为C.UTF-8后问题仍存在。

解决方案

1. 统一用pathlib处理路径(推荐)

Python 3的pathlib模块原生支持非ASCII路径,无需手动编码,避免字符串拼接的编码问题:

写入文件

from pathlib import Path

# 用Path对象拼接路径,自动处理编码
file_path = Path(path) / filename
with file_path.open('wb') as f:
    f.write(binary_data)

读取文件

from pathlib import Path

target_dir = Path(path)
# 遍历目录下的文件
for file_path in target_dir.iterdir():
    # 直接用Path对象打开,自动处理路径编码
    with file_path.open('rb') as f:
        data = f.read()
        # 后续数据处理逻辑

2. 直接使用字节路径打开

如果必须手动处理编码,读取时直接获取路径的字节形式,跳过字符串转码步骤:

from pathlib import Path

target_dir = Path(path)
for file_path in target_dir.iterdir():
    # 用surrogateescape编码获取字节路径
    bytes_path = file_path.as_posix().encode('utf-8', 'surrogateescape')
    with open(bytes_path, 'rb') as f:
        data = f.read()
        # 后续数据处理逻辑

3. 确认Python运行时编码

检查Python的文件系统编码是否为UTF-8:

import sys
print(sys.getfilesystemencoding())

若输出不是utf-8,在运行脚本前设置环境变量:

export LC_ALL=C.UTF-8
export LANG=C.UTF-8

4. 上传环节的文件名处理

确保网站接收上传文件名时,已正确解码为UTF-8字符串,避免保留原始字节串导致路径异常(如Flask/Django框架中,上传文件的filename属性应为已解码的字符串)。

内容的提问来源于stack exchange,提问作者spospider

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 06:06:20