从Feed下载含特殊字符的CSV文件:文件名合规处理方案咨询
解决CSV文件名含非法字符的保存问题
方案1:正则表达式批量替换非法字符
跨平台处理所有操作系统禁止的特殊字符,将其替换为下划线(或空字符串),同时处理空文件名等边界情况:
import re import os def sanitize_filename(filename): # 匹配Windows/Unix类系统的所有非法文件名字符 illegal_pattern = re.compile(r'[\/:*?"<>|]') # 替换非法字符为下划线,去除首尾空白 sanitized = illegal_pattern.sub('_', filename).strip() # 防止处理后文件名为空,设置默认名称 return sanitized if sanitized else "default_csv_file" # 原代码修改后 collectionDownload = 'https://www.myfeedwebsite.com/api/' + collectionID + '/download/csv' response2 = requests.get(collectionDownload, headers=headers) if response2.status_code == 200: original_name = i['attributes']['name'] safe_name = sanitize_filename(original_name) # 使用with语句自动管理文件句柄,比手动close更安全 with open(f"{safe_name}.csv", "wb") as file: file.write(response2.content)
方案2:使用第三方库处理更复杂的文件名场景
如果需要处理特殊字符(如非英文字符、空格转连字符等),可以用python-slugify库,它会自动将文件名转换为安全的格式:
- 先安装库:
pip install python-slugify
- 代码实现:
from slugify import slugify # 原代码修改后 collectionDownload = 'https://www.myfeedwebsite.com/api/' + collectionID + '/download/csv' response2 = requests.get(collectionDownload, headers=headers) if response2.status_code == 200: original_name = i['attributes']['name'] # slugify会自动替换非法字符、处理空格和特殊字符 safe_name = slugify(original_name) with open(f"{safe_name}.csv", "wb") as file: file.write(response2.content)
额外优化:避免文件名重复覆盖
如果存在多个文件名处理后重复的情况,可以在函数中添加重复检测,自动追加序号:
import re import os def sanitize_filename(filename, output_dir="."): illegal_pattern = re.compile(r'[\/:*?"<>|]') sanitized = illegal_pattern.sub('_', filename).strip() sanitized = sanitized if sanitized else "default_csv_file" base_name, _ = os.path.splitext(sanitized) counter = 1 final_name = f"{base_name}.csv" # 检查文件是否存在,存在则追加序号 while os.path.exists(os.path.join(output_dir, final_name)): final_name = f"{base_name}_{counter}.csv" counter += 1 return final_name # 使用示例 safe_name = sanitize_filename(i['attributes']['name']) with open(safe_name, "wb") as file: file.write(response2.content)
内容的提问来源于stack exchange,提问作者Gwynbleidd
相关产品推荐
相关产品推荐

