如何读取文本文件,提取RTE编号并生成关联列?
文本处理需求
原始输入文本
RTE 001 CO. CITY POSTMILE PT POINT ORA DAPT R000.129 DH 000.102 ORA DAPT R000.204 DI ORA DAPT R000.231 DH 000.022 .......... RTE 022 CO. CITY POSTMILE PT POINT ORA SLB 000.000 DH 000.017 ORA SLB 000.017 DH 000.130
期望输出文本
RTE CO. CITY POSTMILE PT POINT 001 ORA DAPT R000.129 DH 000.102 001 ORA DAPT R000.204 DI 001 ORA DAPT R000.231 DH 000.022 022 ORA SLB 000.000 DH 000.017 022 ORA SLB 000.017 DH 000.130
问题说明
目前可识别并处理以ORA开头的行,但无法提取RTE 001、RTE 022这类编号,并将其作为列添加到后续行中,直到遇到下一个RTE标识为止。
解决方案1:使用AWK脚本
AWK适合快速处理文本行,以下脚本可实现需求:
BEGIN { rte = "" } /^[[:space:]]*RTE/ { split($0, parts, /RTE/) rte = parts[2] gsub(/^[[:space:]]+|[[:space:]]+$/, "", rte) next } /^[[:space:]]*ORA/ { gsub(/^[[:space:]]+/, "", $0) printf "%s %s\n", rte, $0 next } /^[[:space:]]*CO\. CITY/ { gsub(/^[[:space:]]+/, "", $0) printf "RTE %s\n", $0 }
使用方式:将脚本保存为process_rte.awk,执行命令:
awk -f process_rte.awk input.txt > output.txt
解决方案2:使用Python脚本
如果需要更灵活的逻辑,可使用Python实现:
current_rte = "" output_lines = [] with open("input.txt", "r") as f: for line in f: stripped_line = line.strip() # 跳过空行和分隔线 if not stripped_line or stripped_line.startswith(".........."): continue # 提取RTE编号 if stripped_line.startswith("RTE"): current_rte = stripped_line.split()[1] # 处理表头行 elif stripped_line.startswith("CO. CITY"): output_lines.append(f"RTE {stripped_line}") # 处理ORA行,添加RTE编号 elif stripped_line.startswith("ORA"): output_lines.append(f"{current_rte} {stripped_line}") # 写入输出文件 with open("output.txt", "w") as f: f.write("\n\n".join(output_lines) + "\n")
使用方式:保存为process_rte.py,执行命令:
python process_rte.py
内容的提问来源于stack exchange,提问作者tinhtu
相关产品推荐
相关产品推荐

