如何基于TSV值重命名文件及拆分含重复表头的TSV文件?
1. Rename Files Based on Values in a TSV File
First, let's assume your TSV has two columns: the original filename (or full path) and the new name you want to assign. For example, your mapping TSV might look like this:
old_report.tsv Q3_sales_data.tsv
raw_data_001.tsv processed_segment_2.tsv
Option 1: Python Script (Flexible for Edge Cases)
This script reads the TSV, checks if files exist before renaming, and skips invalid rows—perfect if you need extra safety:
import os import csv # Update this to your mapping TSV path mapping_tsv = "file_renames.tsv" with open(mapping_tsv, 'r', newline='') as tsv_file: reader = csv.reader(tsv_file, delimiter='\t') for row_num, row in enumerate(reader, 1): if len(row) < 2: print(f"Skipping row {row_num}: Not enough columns") continue old_name, new_name = row[0].strip(), row[1].strip() if os.path.isfile(old_name): os.rename(old_name, new_name) print(f"Renamed: {old_name} → {new_name}") else: print(f"Warning: {old_name} doesn't exist (row {row_num})")
Option 2: Command Line with awk (Fast for Simple Mappings)
If you prefer working in the terminal, this one-liner handles basic renames on Unix-like systems:
awk -F'\t' '{system("mv -n "$1" "$2)}' file_renames.tsv
The -n flag prevents accidental overwrites—remove it only if you’re sure you want to replace existing files with the new names.
2. Split TSV with Repeating Headers into Separate Files
Your input has repeated Position A B C D headers, and you want each header + its associated data as a standalone TSV. Here are two straightforward methods:
Option 1: Python Script (Easy to Customize)
This script tracks when it hits a header, creates a new output file, and writes lines until the next header comes up:
input_file = "your_input.tsv" output_base_name = "segment_" file_index = 1 current_output = None with open(input_file, 'r') as infile: for line in infile: cleaned_line = line.strip() # Match your exact header text (adjust if your header has different spacing) if cleaned_line == "Position A B C D": # Close the previous file if it exists if current_output: current_output.close() # Create a new output file output_path = f"{output_base_name}{file_index}.tsv" current_output = open(output_path, 'w') file_index += 1 # Write the line to the current file (only if we've started a file) if current_output: current_output.write(line) # Close the last open file if current_output: current_output.close() print(f"Successfully split into {file_index - 1} files!")
This will generate segment_1.tsv, segment_2.tsv, etc., each containing one complete header-data block.
Option 2: Command Line with awk (No Python Required)
For a terminal-only solution, this awk command does the job in one line:
awk '/^Position A B C D$/ {out_file="split_"++count".tsv"} {print > out_file}' your_input.tsv
Every time it detects the header line, it creates a new file named split_1.tsv, split_2.tsv, etc., and writes all subsequent lines to that file until the next header is found.
内容的提问来源于stack exchange,提问作者SaltedPork

