You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将大型TXT文件转换为Python字典?

从TXT文件提取术语-定义对并转为Python字典

核心思路

你的TXT文件遵循「术语单独一行,后续一/多行是对应定义,直到下一个术语出现」的结构。我们可以通过判断行的特征区分术语和定义,自动遍历生成字典。

实现代码

def txt_to_terms_dict(file_path):
    terms_dict = {}
    current_term = None
    current_definition = []

    with open(file_path, 'r', encoding='utf-8') as f:
        # 预处理:去除空行,清理每行首尾空白
        lines = [line.strip() for line in f if line.strip()]

    for line in lines:
        # 识别术语:这里以「行末尾无句点」为判断依据,可按需调整
        if not line.endswith('.'):
            # 保存上一组术语-定义对
            if current_term is not None:
                terms_dict[current_term] = ' '.join(current_definition)
                current_definition = []
            current_term = line
        else:
            # 收集定义行
            current_definition.append(line)
    
    # 处理最后一组(避免遗漏末尾术语)
    if current_term is not None:
        terms_dict[current_term] = ' '.join(current_definition) if current_definition else ''

    return terms_dict

# 调用示例
if __name__ == '__main__':
    tech_terms = txt_to_terms_dict('your_terms_file.txt')
    # 验证结果
    for term, desc in tech_terms.items():
        print(f"【术语】{term}\n【定义】{desc}\n")

调整优化

如果部分术语也带句点,可改用正则匹配技术术语特征,比如针对网络术语的正则:

import re

def txt_to_terms_dict(file_path):
    terms_dict = {}
    current_term = None
    current_definition = []
    # 匹配网络术语模式:含数字+BASE/VG/AnyLan,或数字:数字格式
    term_pattern = re.compile(r'^\d+.*(BASE|VG|AnyLan)|^\d+:\d+$', re.IGNORECASE)

    with open(file_path, 'r', encoding='utf-8') as f:
        lines = [line.strip() for line in f if line.strip()]

    for line in lines:
        if term_pattern.match(line):
            if current_term is not None:
                terms_dict[current_term] = ' '.join(current_definition)
                current_definition = []
            current_term = line
        else:
            current_definition.append(line)
    
    if current_term is not None:
        terms_dict[current_term] = ' '.join(current_definition) if current_definition else ''

    return terms_dict

特殊情况处理

  • 无定义的术语(如示例中的16:9)会被赋值为空字符串,若需跳过这类条目,可在最后添加过滤逻辑:
    # 过滤空定义的术语
    terms_dict = {k: v for k, v in terms_dict.items() if v}
    

内容的提问来源于stack exchange,提问作者Xavier Perez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 18:15:41