You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python移除TXT文件前20行注释并导出表格为CSV

用Python移除TXT注释行并转存为CSV

单个文件处理

假设你的TXT文件开头20行是注释,剩余内容为制表符分隔的表格(NCBI dbGap文件通常采用这种格式),可以用以下两种方式处理:

方法1:直接读写(简单场景适用)

# 定义文件路径
input_path = r"C:\Users\test.txt"
output_path = r"C:\Users\output.csv"

# 跳过注释行并转换格式
with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile:
    # 跳过前20行注释
    for _ in range(20):
        next(infile)
    # 将制表符替换为逗号,写入CSV
    for line in infile:
        csv_line = line.replace('\t', ',')
        outfile.write(csv_line)

方法2:用csv模块(规范处理复杂表格)

如果需要处理带引号的字段或不确定分隔符,用csv模块更稳妥:

import csv

input_path = r"C:\Users\test.txt"
output_path = r"C:\Users\output.csv"

with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile:
    # 跳过前20行注释
    for _ in range(20):
        next(infile)
    
    # 按制表符读取原文件,写入CSV格式
    reader = csv.reader(infile, delimiter='\t')
    writer = csv.writer(outfile)
    
    for row in reader:
        writer.writerow(row)

多个文件批量处理

如果有大量同格式TXT文件,用os模块遍历处理:

import csv
import os

# 源文件文件夹和输出文件夹路径
source_folder = r"C:\Users\your_txt_folder"
output_folder = r"C:\Users\csv_output"

# 遍历所有TXT文件
for filename in os.listdir(source_folder):
    if filename.endswith('.txt'):
        input_path = os.path.join(source_folder, filename)
        output_filename = f"{os.path.splitext(filename)[0]}.csv"
        output_path = os.path.join(output_folder, output_filename)
        
        with open(input_path, 'r', encoding='utf-8') as infile, open(output_path, 'w', encoding='utf-8', newline='') as outfile:
            # 跳过前20行注释
            for _ in range(20):
                next(infile)
            
            reader = csv.reader(infile, delimiter='\t')
            writer = csv.writer(outfile)
            
            for row in reader:
                writer.writerow(row)

额外优化:按注释特征跳过行

如果不同文件的注释行数不固定,可通过注释行特征(比如以#开头)判断跳过:

with open(input_path, 'r', encoding='utf-8') as infile:
    # 跳过所有以#开头的行
    for line in infile:
        if not line.strip().startswith('#'):
            # 将文件指针移回表格起始行
            infile.seek(infile.tell() - len(line))
            break
    # 后续读取处理逻辑同上

注意事项

  • 编码问题:若读取出现乱码,可尝试将encoding='utf-8'替换为encoding='latin-1'(NCBI文件常用该编码)。

内容的提问来源于stack exchange,提问作者Jim_13

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 20:04:54