如何修改load_data函数区分同后缀的type3与type4 TXT文件?
解决load_data函数中type3与type4 TXT文件混淆的方案
原函数存在两个明显问题:一是type0分支里重复判断了.txt后缀,导致type4的分支永远无法触发;二是type3和type4同属TXT后缀,且要求不能打开文件区分,只能从用户显式指定或文件名约定两个方向解决。
方案一:强制用户显式指定格式
既然自动识别无法区分两类TXT文件,干脆在type0模式下遇到TXT文件时直接提示用户手动指定type3或type4,彻底避免混淆:
from typing import Literal def load_data(path: str, format: Literal['type0', 'type1', 'type2', 'type3', 'type4']) -> Data: valid_formats = {'type0', 'type1', 'type2', 'type3', 'type4'} if format not in valid_formats: raise ValueError(f"Incorrect format `{format}` specified for load_data()") if format == 'type0': if path.lower().endswith('.pdf'): data = parse_type1(path) elif path.lower().endswith('.csv') or path.lower().endswith('.d'): data = parse_type2(path) elif path.lower().endswith('.txt'): # 无法自动区分TXT类型,提示用户显式指定 raise ValueError("TXT files can be either type3 (comma-separated) or type4 (space-separated). " "Please specify the format explicitly instead of using 'type0'.") else: raise Exception("Unknown file format in load_data(), consider specifying the format instead of using `type0`") elif format == 'type1': data = parse_type1(path) elif format == 'type2': data = parse_type2(path) elif format == 'type3': data = parse_type3(path) elif format == 'type4': data = parse_type4(path) # 校验数据结构一致性 assert data.data.shape[0] == data.wavelength.shape[0], 'Parsing raw data by load_data() yields inconsistent shapes' assert data.data.shape[1] == data.time.shape[0], 'Parsing raw data by load_data() yields inconsistent shapes' return data
方案二:通过文件名约定区分
如果希望保留type0的自动识别能力,可以和团队约定文件名规则:比如type3的TXT文件名包含_comma,type4的包含_space,通过文件名关键字判断:
from typing import Literal def load_data(path: str, format: Literal['type0', 'type1', 'type2', 'type3', 'type4']) -> Data: valid_formats = {'type0', 'type1', 'type2', 'type3', 'type4'} if format not in valid_formats: raise ValueError(f"Incorrect format `{format}` specified for load_data()") if format == 'type0': if path.lower().endswith('.pdf'): data = parse_type1(path) elif path.lower().endswith('.csv') or path.lower().endswith('.d'): data = parse_type2(path) elif path.lower().endswith('.txt'): filename = path.lower() # 根据约定的文件名关键字区分 if '_comma' in filename: data = parse_type3(path) elif '_space' in filename: data = parse_type4(path) else: raise ValueError("TXT files need to include '_comma' (for type3) or '_space' (for type4) in filename " "when using 'type0', or specify the format explicitly.") else: raise Exception("Unknown file format in load_data(), consider specifying the format instead of using `type0`") elif format == 'type1': data = parse_type1(path) elif format == 'type2': data = parse_type2(path) elif format == 'type3': data = parse_type3(path) elif format == 'type4': data = parse_type4(path) assert data.data.shape[0] == data.wavelength.shape[0], 'Parsing raw data by load_data() yields inconsistent shapes' assert data.data.shape[1] == data.time.shape[0], 'Parsing raw data by load_data() yields inconsistent shapes' return data
说明
两个方案都满足"无需打开文件"的要求:
- 方案一简单直接,完全依赖用户明确声明格式,避免了任何自动识别的歧义;
- 方案二保留了自动识别的便利性,但需要团队严格遵守文件名约定。
内容的提问来源于stack exchange,提问作者jasonstich
相关产品推荐
相关产品推荐

