You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于TSV值重命名文件及拆分含重复表头的TSV文件?

Two Solutions for Your TSV Processing Tasks

1. Rename Files Based on Values in a TSV File

First, let's assume your TSV has two columns: the original filename (or full path) and the new name you want to assign. For example, your mapping TSV might look like this:

old_report.tsv Q3_sales_data.tsv
raw_data_001.tsv processed_segment_2.tsv

Option 1: Python Script (Flexible for Edge Cases)

This script reads the TSV, checks if files exist before renaming, and skips invalid rows—perfect if you need extra safety:

import os
import csv

# Update this to your mapping TSV path
mapping_tsv = "file_renames.tsv"

with open(mapping_tsv, 'r', newline='') as tsv_file:
    reader = csv.reader(tsv_file, delimiter='\t')
    for row_num, row in enumerate(reader, 1):
        if len(row) < 2:
            print(f"Skipping row {row_num}: Not enough columns")
            continue
        old_name, new_name = row[0].strip(), row[1].strip()
        if os.path.isfile(old_name):
            os.rename(old_name, new_name)
            print(f"Renamed: {old_name} → {new_name}")
        else:
            print(f"Warning: {old_name} doesn't exist (row {row_num})")

Option 2: Command Line with awk (Fast for Simple Mappings)

If you prefer working in the terminal, this one-liner handles basic renames on Unix-like systems:

awk -F'\t' '{system("mv -n "$1" "$2)}' file_renames.tsv

The -n flag prevents accidental overwrites—remove it only if you’re sure you want to replace existing files with the new names.


2. Split TSV with Repeating Headers into Separate Files

Your input has repeated Position A B C D headers, and you want each header + its associated data as a standalone TSV. Here are two straightforward methods:

Option 1: Python Script (Easy to Customize)

This script tracks when it hits a header, creates a new output file, and writes lines until the next header comes up:

input_file = "your_input.tsv"
output_base_name = "segment_"
file_index = 1
current_output = None

with open(input_file, 'r') as infile:
    for line in infile:
        cleaned_line = line.strip()
        # Match your exact header text (adjust if your header has different spacing)
        if cleaned_line == "Position A B C D":
            # Close the previous file if it exists
            if current_output:
                current_output.close()
            # Create a new output file
            output_path = f"{output_base_name}{file_index}.tsv"
            current_output = open(output_path, 'w')
            file_index += 1
        # Write the line to the current file (only if we've started a file)
        if current_output:
            current_output.write(line)

# Close the last open file
if current_output:
    current_output.close()

print(f"Successfully split into {file_index - 1} files!")

This will generate segment_1.tsv, segment_2.tsv, etc., each containing one complete header-data block.

Option 2: Command Line with awk (No Python Required)

For a terminal-only solution, this awk command does the job in one line:

awk '/^Position A B C D$/ {out_file="split_"++count".tsv"} {print > out_file}' your_input.tsv

Every time it detects the header line, it creates a new file named split_1.tsv, split_2.tsv, etc., and writes all subsequent lines to that file until the next header is found.


内容的提问来源于stack exchange,提问作者SaltedPork

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:09:21