You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取包含两个堆叠数据集的固定宽度数据文件?

Reading a Fixed-Width File with Two Separate Datasets

I’ve dealt with this exact kind of file structure before—split datasets in a single fixed-width file can be tricky, but using Python’s pandas library makes it straightforward. Here’s how to handle it, plus a command-line alternative if you prefer that.

Using Python & Pandas

First, make sure you have pandas installed (pip install pandas if you don’t). Then, follow these steps:

  1. Define your column widths for each dataset. For example, if Dataset 1 has 3 columns with widths 10, 15, and 8 characters, and Dataset 2 has 3 columns with widths 5, 20, and 10, you’d set:

    dataset1_widths = [10, 15, 8]
    dataset2_widths = [5, 20, 10]
    
  2. Read Dataset 1: This includes the header (line 1) and data lines 2-1500 (total 1499 data rows). We’ll use pd.read_fwf() which is designed for fixed-width files:

    import pandas as pd
    
    # Read Dataset 1
    df1 = pd.read_fwf(
        "your_file.txt",
        widths=dataset1_widths,
        header=0,  # Use line 1 as the header
        nrows=1499,  # Read lines 2 to 1500 (1499 rows total)
        skiprows=0  # Don't skip any lines at the start
    )
    
  3. Read Dataset 2: This starts at line 1501 (the header) and includes data lines 1502-3001 (another 1499 rows). We need to skip the first 1500 lines to get to its header:

    # Read Dataset 2
    df2 = pd.read_fwf(
        "your_file.txt",
        widths=dataset2_widths,
        header=0,  # Use line 1501 as the header
        nrows=1499,  # Read lines 1502 to 3001
        skiprows=1500  # Skip first 1500 lines (Dataset 1's header + data)
    )
    

That’s it! You now have two DataFrames (df1 and df2) with each dataset’s data properly parsed. Adjust the widths, nrows, and skiprows values to match your actual file structure.

Command-Line Alternative (Using Awk)

If you prefer working in the terminal, you can use awk to extract and parse each dataset. Let’s use the same width examples as above:

  • Extract Dataset 1:

    awk 'NR >=2 && NR <=1500 {print substr($0,1,10), substr($0,11,15), substr($0,26,8)}' your_file.txt > dataset1.csv
    # Add the header line separately:
    head -n1 your_file.txt | awk '{print substr($0,1,10), substr($0,11,15), substr($0,26,8)}' > temp.csv && cat temp.csv dataset1.csv > dataset1_final.csv && rm temp.csv
    
  • Extract Dataset 2:

    awk 'NR >=1502 && NR <=3001 {print substr($0,1,5), substr($0,6,20), substr($0,26,10)}' your_file.txt > dataset2.csv
    # Add the header line:
    sed -n '1501p' your_file.txt | awk '{print substr($0,1,5), substr($0,6,20), substr($0,26,10)}' > temp.csv && cat temp.csv dataset2.csv > dataset2_final.csv && rm temp.csv
    

In the substr commands, the syntax is substr(string, start_position, length). Adjust these numbers to match your actual column widths.


内容的提问来源于stack exchange,提问作者Union find

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:06:49