如何使用Python依据指定文件的内容顺序对一组文件进行排序?
Got it, let's solve this problem where you need to reorder the rows in file_B.txt and file_C.txt exactly to match the order in file_A.txt—while keeping all original values intact. Below are two straightforward approaches, depending on which tools you prefer working with.
Approach 1: Using Shell Tools (awk + bash)
This is perfect if you're comfortable with command-line workflows. The core idea is to first capture the exact ID sequence from file_A.txt, then rearrange the rows in target files to follow that sequence.
Run these commands in your terminal:
# Step 1: Extract header and ID order from the reference file (file_A.txt) head -n 1 file_A.txt > temp_header.txt awk 'NR>1 {print $1}' file_A.txt > id_order.txt # Step 2: Reorder file_B.txt # Map each ID to its full line in file_B.txt awk 'NR==1{next} {map[$1]=$0}' file_B.txt > temp_B_map.txt # Rebuild file_B.txt using the reference ID order cat temp_header.txt > file_B.txt while read id; do grep "^$id " temp_B_map.txt >> file_B.txt done < id_order.txt # Step 3: Repeat the process for file_C.txt awk 'NR==1{next} {map[$1]=$0}' file_C.txt > temp_C_map.txt cat temp_header.txt > file_C.txt while read id; do grep "^$id " temp_C_map.txt >> file_C.txt done < id_order.txt # Clean up temporary files rm temp_header.txt id_order.txt temp_B_map.txt temp_C_map.txt
How this works:
- We first save the shared header and the precise order of IDs from
file_A.txt. - For each target file, we create a lookup of IDs to their full line content.
- We then rebuild the target file by writing the header first, followed by lines ordered to match the reference ID sequence.
Approach 2: Using Python Script
If you prefer a more readable, maintainable solution, Python is a great choice. This script handles the alignment with minimal manual steps.
Create a file named align_files.py with this code:
def align_to_reference(reference_path, target_path): # Read reference file to get header and ID sequence with open(reference_path, 'r') as ref_file: lines = ref_file.readlines() header = lines[0] # Extract IDs from non-header lines, preserving their order id_sequence = [line.split()[0] for line in lines[1:] if line.strip()] # Map each ID in target file to its full line content id_to_line = {} with open(target_path, 'r') as target_file: # Skip the target's header (we'll use the reference's header instead) next(target_file) for line in target_file: line = line.strip() if not line: continue id_val = line.split()[0] id_to_line[id_val] = line # Write aligned content back to the target file with open(target_path, 'w') as target_file: target_file.write(header) for id_val in id_sequence: target_file.write(f"{id_to_line[id_val]}\n") # Align both files to match file_A.txt align_to_reference('file_A.txt', 'file_B.txt') align_to_reference('file_A.txt', 'file_C.txt')
Run the script with:
python align_files.py
Notes:
- Both approaches assume every ID in
file_A.txtexists infile_B.txtandfile_C.txt(matching your example). If there are missing IDs, you can add logic to append those missing lines at the end of target files. - If your files have duplicate IDs, adjust the logic to handle multiple entries per ID (e.g., collecting all lines for an ID instead of overwriting).
内容的提问来源于stack exchange,提问作者Iacopo Passeri

