You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从TPM手册文本提取机器缺陷与方案并构建指定结构字典?

问题与需求

我有一个包含4000行的TPM手册文本文件,格式固定:

  • 首行是机器标题
  • 后续内容为缺陷条目(以Defect或Defeito开头),每个缺陷条目下方是对应的解决方案列表

文件内容示例:

Balancim de corte hidráulico (a) ponte
Defect 01 – Máquina não liga
Botão de emergência acionado
Problema no pedal
Defeito 02 – O martelo não vai para os lados
Botão de emergência acionado
...

需要生成如下结构的字典:

machine_dict = {
    'Balancim de corte hidráulico (a) ponte': {
        'Defect 01 – Máquina não liga': ['Botão de emergência acionado', 'Problema no pedal', ...],
        'Defeito 02 – O martelo não vai para os lados': ['Botão de emergência acionado', ...]
    }
}

当前代码无法正确分离缺陷与解决方案,代码如下:

with open("Manual TPM/manual.txt") as manual:
    with open('Manual TPM/manual.txt') as manual:
        manual_tpm = manual.read()  

        maqs_defeito = [list.split('\n') for list in manual_tpm.split('\n\n') if list]

        maqs = {}
        defeitos=[]
        solucoes=[]

        for index, item in enumerate(maqs_defeito):
            maqs[item[0]] = {}
            if "Defeito" in item:
                defeitos.append([item,index])
            solucoes.append(item)

        for i in range(len(defeitos)):
            maqs[item[0]][defeitos[i][0]] = solucoes[defeitos[i][1]:]
        print(maqs)

修正后的代码
machine_dict = {}

with open("Manual TPM/manual.txt", encoding="utf-8") as f:
    # 过滤空行并去除每行首尾空格,避免无效内容干扰
    lines = [line.strip() for line in f if line.strip()]

current_machine = None
current_defect = None

for line in lines:
    # 判定机器标题:既不是Defect也不是Defeito开头的行
    if not line.startswith(("Defect", "Defeito")):
        current_machine = line
        machine_dict[current_machine] = {}
        continue
    
    # 判定缺陷条目:以Defect/Defeito开头的行
    if line.startswith(("Defect", "Defeito")):
        current_defect = line
        machine_dict[current_machine][current_defect] = []
        continue
    
    # 剩余行作为解决方案,添加到当前缺陷的列表中
    machine_dict[current_machine][current_defect].append(line)

print(machine_dict)

关键修正说明
  1. 移除冗余操作:删除原代码中重复打开文件的无效逻辑
  2. 精准逐行匹配:放弃按空行分割的模糊方式,改为逐行遍历,严格匹配格式规则:
    • 非缺陷开头的行判定为机器标题,初始化对应字典
    • 缺陷开头的行判定为缺陷条目,初始化解决方案列表
    • 其余行直接作为解决方案加入对应列表
  3. 预处理内容:提前过滤空行和行首尾空格,避免无效内容干扰
  4. 编码兼容:添加encoding="utf-8"参数,适配葡萄牙语等非ASCII字符的编码需求

内容的提问来源于stack exchange,提问作者user23059245

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 15:07:46