如何对TXT文件按指定列排序另存?含Linux命令及Python实现
Got it, let's break down your text file sorting needs with both Linux terminal commands and a Python script using NumPy—right up your alley!
Linux Terminal Approach
You already have the core sort command down, so let's just fix the output saving part first.
Sort by Column 3 & Save to New File
To save your sorted results instead of just printing them to the screen, you have two straightforward options:
Redirect output with
>: This sends the sorted text straight to a new file.sort -k3n myfile.txt > sorted_by_col3.txtNote: If
sorted_by_col3.txtalready exists, this will overwrite it. If you ever need to append instead (not necessary here), use>>instead of>.Use the
-oflag: This is the dedicated output flag forsort, and it's safer if you're worried about accidentally overwriting files.sort -k3n myfile.txt -o sorted_by_col3.txtBonus: This even works if you wanted to sort a file in-place (though I'd always recommend saving to a new file to keep your original data intact!).
Sort by Column 4 (Your Second Terminal Request)
Switching to sort by the 4th column is just a quick tweak to the k parameter—change it to 4:
sort -k4n myfile.txt > sorted_by_col4.txt
Remember the n flag is crucial here: it tells sort to do numerical sorting instead of lexicographical (which would mess up numbers like "91.49" vs "40.65").
Python Script with NumPy (Load → Sort → Save → Re-read)
If you want to do this programmatically with Python and numpy.loadtxt, here's a complete, easy-to-follow script that handles every step you asked for—including preserving the header from your sample data.
Full Working Script
import numpy as np # Step 1: Read the original file, keeping the header separate with open("myfile.txt", "r") as input_file: # Grab the first line as the header header_line = input_file.readline().strip() # Load the rest of the data into a numpy array (skip the header row) raw_data = np.loadtxt("myfile.txt", skiprows=1) # Step 2: Sort the data by column 3 (note: NumPy uses 0-indexing, so column 3 = index 2) # To sort by column 4 instead, change the index to 3 sorted_data = raw_data[raw_data[:, 2].argsort()] # Step 3: Save the sorted data to a new file, including the header with open("sorted_by_col3_numpy.txt", "w") as output_file: output_file.write(header_line + "\n") # Use %g to match your sample data's formatting (auto scientific/fixed notation) np.savetxt(output_file, sorted_data, fmt="%g") # Step 4: Read the new sorted file back to verify it worked with open("sorted_by_col3_numpy.txt", "r") as sorted_file: sorted_header = sorted_file.readline().strip() loaded_sorted_data = np.loadtxt("sorted_by_col3_numpy.txt", skiprows=1) # Quick check to confirm everything went right print(f"Original data had {raw_data.shape[0]} rows") print(f"Sorted data loaded has {loaded_sorted_data.shape[0]} rows") print("First 3 rows of sorted data:\n", loaded_sorted_data[:3])
Quick Notes to Customize:
- Column Indexing: NumPy counts columns starting at 0. So if you want to sort by column 4 instead of 3, just change
raw_data[:, 2]toraw_data[:, 3]. - Header Handling: We read the header first because
numpy.loadtxtwould treat it as numerical data otherwise. Writing it back keeps your output file structured just like the original. - Formatting: The
fmt="%g"ensures numbers are saved in a clean format, matching the scientific notation and decimals from your sample data.
Sample Data Reference
Here's the sample data you provided, for easy reference:
# ID x y 6.60968219252e-05 7.56508909895e-10 40.65 0.000196142774512 1.90541971372e-09 49.18 0.000451120770124 3.75884511195e-09 60.78 9.49736290045e-05 1.08754058315e-09 44.12 0.000773197066156 5.55965157568e-09 70.64 0.000395119768811 5.35886928694e-09 48.42 0.000761797071911 1.1411313874e-08 42.8 6.13793543105e-05 6.79135943796e-10 36.94 0.0014257833689 6.69702707603e-09 91.49 8.02798012773e-05 8.34778117262e-10 43.19
内容的提问来源于stack exchange,提问作者bhjghjh

