Python 3如何从.txt格式的表格中提取指定列数据并生成输出文件
Python3提取txt表格指定列方案
核心逻辑
不管是制表符分隔的纯文本表格还是竖线分隔的Markdown表格,本质都是按固定分隔符拆分的结构化数据,只需要先定位目标列的位置,再逐行提取即可,无需额外安装第三方依赖。
对应实际制表符分隔表格的代码
你贴的实际数据为制表符分隔格式,可直接使用以下代码:
import csv # 可根据实际需求修改以下配置项 input_path = "你的输入文件路径.txt" output_path = "输出结果.txt" # 要提取的列名,提取其他列直接修改这里即可 target_cols = ["Input", "Surname"] with open(input_path, 'r', encoding='utf-8') as in_file, open(output_path, 'w', encoding='utf-8', newline='') as out_file: reader = csv.DictReader(in_file, delimiter='\t') writer = csv.DictWriter(out_file, fieldnames=target_cols, delimiter='\t') # 写入表头 writer.writeheader() for row in reader: # 如果你需要仅保留Input为GO:0003723开头的行,放开下面两行注释即可 # if not row['Input'].startswith('GO:0003723'): # continue writer.writerow({col: row[col] for col in target_cols})
对应示例Markdown格式表格的代码
如果你要处理的是竖线分隔的Markdown格式表格,使用以下代码:
import csv input_path = "你的输入文件路径.txt" output_path = "输出结果.txt" target_cols = ["Input", "Surname"] with open(input_path, 'r', encoding='utf-8') as in_file, open(output_path, 'w', encoding='utf-8', newline='') as out_file: # 处理表头,去除多余空格和空值 header = [col.strip() for col in next(in_file).split('|') if col.strip()] # 跳过Markdown表格的分隔线行(即|---|...|那一行) next(in_file) writer = csv.DictWriter(out_file, fieldnames=target_cols, delimiter='\t') writer.writeheader() for line in in_file: line = line.strip() if not line: continue row_data = [col.strip() for col in line.split('|') if col.strip()] row = dict(zip(header, row_data)) writer.writerow({col: row[col] for col in target_cols})
注意事项
- 运行前请将
input_path和output_path替换为你本地的实际文件路径 - 如果实际需求提取的列名和示例不同,直接修改
target_cols的内容即可
内容的提问来源于stack exchange,提问作者Pythonstudent
相关产品推荐
相关产品推荐

