如何将大型TXT文件转换为Python字典?
从TXT文件提取术语-定义对并转为Python字典
核心思路
你的TXT文件遵循「术语单独一行,后续一/多行是对应定义,直到下一个术语出现」的结构。我们可以通过判断行的特征区分术语和定义,自动遍历生成字典。
实现代码
def txt_to_terms_dict(file_path): terms_dict = {} current_term = None current_definition = [] with open(file_path, 'r', encoding='utf-8') as f: # 预处理:去除空行,清理每行首尾空白 lines = [line.strip() for line in f if line.strip()] for line in lines: # 识别术语:这里以「行末尾无句点」为判断依据,可按需调整 if not line.endswith('.'): # 保存上一组术语-定义对 if current_term is not None: terms_dict[current_term] = ' '.join(current_definition) current_definition = [] current_term = line else: # 收集定义行 current_definition.append(line) # 处理最后一组(避免遗漏末尾术语) if current_term is not None: terms_dict[current_term] = ' '.join(current_definition) if current_definition else '' return terms_dict # 调用示例 if __name__ == '__main__': tech_terms = txt_to_terms_dict('your_terms_file.txt') # 验证结果 for term, desc in tech_terms.items(): print(f"【术语】{term}\n【定义】{desc}\n")
调整优化
如果部分术语也带句点,可改用正则匹配技术术语特征,比如针对网络术语的正则:
import re def txt_to_terms_dict(file_path): terms_dict = {} current_term = None current_definition = [] # 匹配网络术语模式:含数字+BASE/VG/AnyLan,或数字:数字格式 term_pattern = re.compile(r'^\d+.*(BASE|VG|AnyLan)|^\d+:\d+$', re.IGNORECASE) with open(file_path, 'r', encoding='utf-8') as f: lines = [line.strip() for line in f if line.strip()] for line in lines: if term_pattern.match(line): if current_term is not None: terms_dict[current_term] = ' '.join(current_definition) current_definition = [] current_term = line else: current_definition.append(line) if current_term is not None: terms_dict[current_term] = ' '.join(current_definition) if current_definition else '' return terms_dict
特殊情况处理
- 无定义的术语(如示例中的
16:9)会被赋值为空字符串,若需跳过这类条目,可在最后添加过滤逻辑:# 过滤空定义的术语 terms_dict = {k: v for k, v in terms_dict.items() if v}
内容的提问来源于stack exchange,提问作者Xavier Perez
相关产品推荐
相关产品推荐

