如何读取包含两个堆叠数据集的固定宽度数据文件?
I’ve dealt with this exact kind of file structure before—split datasets in a single fixed-width file can be tricky, but using Python’s pandas library makes it straightforward. Here’s how to handle it, plus a command-line alternative if you prefer that.
Using Python & Pandas
First, make sure you have pandas installed (pip install pandas if you don’t). Then, follow these steps:
Define your column widths for each dataset. For example, if Dataset 1 has 3 columns with widths 10, 15, and 8 characters, and Dataset 2 has 3 columns with widths 5, 20, and 10, you’d set:
dataset1_widths = [10, 15, 8] dataset2_widths = [5, 20, 10]Read Dataset 1: This includes the header (line 1) and data lines 2-1500 (total 1499 data rows). We’ll use
pd.read_fwf()which is designed for fixed-width files:import pandas as pd # Read Dataset 1 df1 = pd.read_fwf( "your_file.txt", widths=dataset1_widths, header=0, # Use line 1 as the header nrows=1499, # Read lines 2 to 1500 (1499 rows total) skiprows=0 # Don't skip any lines at the start )Read Dataset 2: This starts at line 1501 (the header) and includes data lines 1502-3001 (another 1499 rows). We need to skip the first 1500 lines to get to its header:
# Read Dataset 2 df2 = pd.read_fwf( "your_file.txt", widths=dataset2_widths, header=0, # Use line 1501 as the header nrows=1499, # Read lines 1502 to 3001 skiprows=1500 # Skip first 1500 lines (Dataset 1's header + data) )
That’s it! You now have two DataFrames (df1 and df2) with each dataset’s data properly parsed. Adjust the widths, nrows, and skiprows values to match your actual file structure.
Command-Line Alternative (Using Awk)
If you prefer working in the terminal, you can use awk to extract and parse each dataset. Let’s use the same width examples as above:
Extract Dataset 1:
awk 'NR >=2 && NR <=1500 {print substr($0,1,10), substr($0,11,15), substr($0,26,8)}' your_file.txt > dataset1.csv # Add the header line separately: head -n1 your_file.txt | awk '{print substr($0,1,10), substr($0,11,15), substr($0,26,8)}' > temp.csv && cat temp.csv dataset1.csv > dataset1_final.csv && rm temp.csvExtract Dataset 2:
awk 'NR >=1502 && NR <=3001 {print substr($0,1,5), substr($0,6,20), substr($0,26,10)}' your_file.txt > dataset2.csv # Add the header line: sed -n '1501p' your_file.txt | awk '{print substr($0,1,5), substr($0,6,20), substr($0,26,10)}' > temp.csv && cat temp.csv dataset2.csv > dataset2_final.csv && rm temp.csv
In the substr commands, the syntax is substr(string, start_position, length). Adjust these numbers to match your actual column widths.
内容的提问来源于stack exchange,提问作者Union find

