从制表符分隔文件移除指定字段/列的技术实现需求
Got it, let's tackle these two common tab-separated value (TSV) editing tasks using awk—it's the go-to tool for this kind of field-level wrangling, and it keeps things clean and efficient.
Case 1: Remove All Fields Containing a Specific String
Goal
Take a TSV file, and for every line, remove any field that includes your target string (e.g., "four"). Output the remaining fields as a valid single-tab-separated TSV.
Solution Command
awk -F'\t' -v OFS='\t' -v target="four" ' { out = "" for (i = 1; i <= NF; i++) { # Skip fields that contain the target string (use $i != target for exact matches) if ($i !~ target) { out = out OFS $i } } # Remove leading tab added by the first OFS sub(/^\t/, "", out) print out }' your_input.tsv
Example
Input line:
one two three four five six four
Output line:
one two three five six
Breakdown
-F'\t'tells awk to split input lines by tabs.-v OFS='\t'sets the output field separator to a single tab, so we don't get messy spacing.-v target="four"passes your target string into the script (easy to change for different values).- The loop checks each field: if it doesn't contain the target, we add it to our output string.
sub(/^\t/, "", out)cleans up the leading tab that gets added when we start building the output.
Case 2: Remove Specific Columns (Avoid Multiple Tabs)
Goal
Remove columns by their numeric position (e.g., column 4 and 7) from a TSV, ensuring the output stays properly formatted with only single tabs between fields (no empty tab gaps from deleted columns).
Solution Command
awk -F'\t' -v OFS='\t' -v cols_to_remove="4,7" ' { split(cols_to_remove, del_cols, ",") # Create a lookup for columns to delete delete del_map for (i in del_cols) { del_map[del_cols[i]] = 1 } out = "" for (i = 1; i <= NF; i++) { # Keep columns not in the delete list if (!(i in del_map)) { out = out OFS $i } } sub(/^\t/, "", out) print out }' your_input.tsv
Example
Input line:
one two three four five six four
Output line (removing columns 4 and 7):
one two three five six
Breakdown
split(cols_to_remove, del_cols, ",")turns your comma-separated column list into an array.del_mapacts as a quick lookup to check if a column should be deleted.- The loop builds the output by only including columns not in
del_map, so we never create empty tab gaps in the first place. - Again,
sub(/^\t/, "", out)cleans up the leading tab from the output string.
Quick One-Liner for Single Column Removal
If you only need to delete one column (e.g., column 4), you can simplify things:
awk -F'\t' -v OFS='\t' '{for(i=4;i<NF;i++) $i=$(i+1); NF--; print}' your_input.tsv
This shifts all columns after the deleted one left by one, then reduces the field count (NF--) to drop the now-duplicate last field—no extra tabs to clean up!
内容的提问来源于stack exchange,提问作者Chris

