如何使用sed或Python重写文件最后一列?GFF文件格式异常求助
Hey there! Let's fix that misbehaving GFF file where extra tabs in the last column are messing up the format. I've got two solid solutions for you—one using sed for quick command-line fixes, and another using Python for more flexibility.
sed to Fix the File The core idea here is to preserve the first 8 tab-separated columns (which are correctly formatted) and replace any extra tabs in the 9th column with spaces (since GFF's 9th column uses semicolons and spaces for key-value pairs anyway).
Run this command in your terminal:
sed -E ':a; s/(([^\t]+\t){8})\t/\1 /; ta' input.gff > output.gff
Breakdown of the command:
:acreates a loop label so we can keep processing until all extra tabs in the 9th column are replaced.(([^\t]+\t){8})matches the first 8 columns (each column is non-tab characters followed by a tab).s/...\t/\1 /replaces the first extra tab after the 8th column with a space.tajumps back to thealabel to repeat the replacement if another tab is found in the 9th column.
If you prefer a scriptable solution (or need to handle edge cases the sed command might miss), Python is perfect. This script will ensure each line has exactly 9 columns by merging any extra split parts back into the 9th column.
with open("input.gff", "r") as infile, open("output.gff", "w") as outfile: for line in infile: # Strip the newline character but keep other whitespace intact stripped_line = line.rstrip("\n") # Split the line by tabs columns = stripped_line.split("\t") if len(columns) > 9: # Keep the first 8 columns, merge the rest into the 9th fixed_columns = columns[:8] + ["\t".join(columns[8:])] else: fixed_columns = columns # Join back with tabs and write to the output file outfile.write("\t".join(fixed_columns) + "\n")
Customization tip:
If you want to replace those extra tabs in the 9th column with spaces instead of keeping them, just change "\t".join(columns[8:]) to " ".join(columns[8:]).
内容的提问来源于stack exchange,提问作者Kay NewEdge Daramola

