如何用Python移除TXT文件前20行注释并导出表格为CSV
用Python移除TXT注释行并转存为CSV
单个文件处理
假设你的TXT文件开头20行是注释,剩余内容为制表符分隔的表格(NCBI dbGap文件通常采用这种格式),可以用以下两种方式处理:
方法1:直接读写(简单场景适用)
# 定义文件路径 input_path = r"C:\Users\test.txt" output_path = r"C:\Users\output.csv" # 跳过注释行并转换格式 with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile: # 跳过前20行注释 for _ in range(20): next(infile) # 将制表符替换为逗号,写入CSV for line in infile: csv_line = line.replace('\t', ',') outfile.write(csv_line)
方法2:用csv模块(规范处理复杂表格)
如果需要处理带引号的字段或不确定分隔符,用csv模块更稳妥:
import csv input_path = r"C:\Users\test.txt" output_path = r"C:\Users\output.csv" with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile: # 跳过前20行注释 for _ in range(20): next(infile) # 按制表符读取原文件,写入CSV格式 reader = csv.reader(infile, delimiter='\t') writer = csv.writer(outfile) for row in reader: writer.writerow(row)
多个文件批量处理
如果有大量同格式TXT文件,用os模块遍历处理:
import csv import os # 源文件文件夹和输出文件夹路径 source_folder = r"C:\Users\your_txt_folder" output_folder = r"C:\Users\csv_output" # 遍历所有TXT文件 for filename in os.listdir(source_folder): if filename.endswith('.txt'): input_path = os.path.join(source_folder, filename) output_filename = f"{os.path.splitext(filename)[0]}.csv" output_path = os.path.join(output_folder, output_filename) with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile: # 跳过前20行注释 for _ in range(20): next(infile) reader = csv.reader(infile, delimiter='\t') writer = csv.writer(outfile) for row in reader: writer.writerow(row)
额外优化:按注释特征跳过行
如果不同文件的注释行数不固定,可通过注释行特征(比如以#开头)判断跳过:
with open(input_path, 'r', encoding='utf-8') as infile: # 跳过所有以#开头的行 for line in infile: if not line.strip().startswith('#'): # 将文件指针移回表格起始行 infile.seek(infile.tell() - len(line)) break # 后续读取处理逻辑同上
注意事项
- 编码问题:若读取出现乱码,可尝试将
encoding='utf-8'替换为encoding='latin-1'(NCBI文件常用该编码)。
内容的提问来源于stack exchange,提问作者Jim_13
相关产品推荐
相关产品推荐

