合并多文本字段并保留内容及字段顺序的Shell脚本求助
Hey Jim, I totally get why join was giving you headaches here—its core design is built around sorted keys and pairwise file comparisons, which makes it tricky to preserve the original order of fields when merging multiple files at once. Let's ditch the join chains and use awk instead, which is way more flexible for this kind of task.
The Problem with join
GNU join -a works great for keeping unpaired lines when combining two files, but when you chain it across multiple files, it enforces sorting on the key field each time. Plus, it doesn’t track the original order that fields first appeared in each file, which is exactly what you need to keep intact.
Solution: Use awk to Track Order and Merge
This awk script will:
- Keep track of the first occurrence order of each key (since you don’t care about key sorting, just field order)
- Append all fields from every file to their corresponding key, in the order they appear
- Output everything in the original key appearance order
Here’s the script:
awk ' BEGIN { # Set your input/output field separators here (adjust to match your files) FS = "\t" # Input separator (e.g., "\t" for tabs, "," for CSV) OFS = "\t" # Output separator (match your desired final format) } { # Grab the first field as the key key = $1 # Collect all fields after the first one (preserve their original order) fields = "" for (i=2; i<=NF; i++) { fields = fields (fields ? OFS : "") $i } # Record the order of first occurrence for each key if (!(key in key_sequence)) { key_sequence[key] = ++sequence_count } # Append the current fields to the key's accumulated list if (key in merged_fields) { merged_fields[key] = merged_fields[key] OFS fields } else { merged_fields[key] = fields } } END { # Sort keys by their first occurrence order and print PROCINFO["sorted_in"] = "@val_num_asc" for (key in key_sequence) { print key, merged_fields[key] } } ' file1.txt file2.txt file3.txt ... # Add all your files here
How to Customize It
- Field Separators: Change
FSandOFSto match your file format (e.g.,FS = ","for CSV files). - Field Delimiter Between Files: If you want a different separator between fields from different files (instead of using
OFS), replace theOFSinmerged_fields[key] = merged_fields[key] OFS fieldswith something else (like";").
Example
Suppose you have two files:file1.txt:
apple red round banana yellow long
file2.txt:
apple sweet crunchy cherry red small
Running the script with these files will output:
apple red round sweet crunchy banana yellow long cherry red small
Notice how apple’s fields keep the order from file1 first, then file2, and the keys are ordered by their first appearance (apple, banana, cherry).
This approach avoids the sorting constraints of join and ensures your field order stays exactly as it appeared in the original files.
内容的提问来源于stack exchange,提问作者Jim Tuck

