You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SED/AWK脚本从超大规模学生文件中提取联系信息

Efficiently Extract Contact Numbers from Large Fixed-Width Student File

Got it, let's tackle this problem head-on—processing a 1M+ line fixed-width file needs to be memory-efficient and fast, since loading the entire file into RAM could cause issues. Here are two robust approaches tailored for your use case:

Key Field Positions (Clarified)

First, let's map your fixed-width specs to code-friendly indices (note: different tools use 0 or 1-based indexing):

  • Student ID: 1-10 characters
  • Mobile Phone: 12-21 characters
  • Home Phone: 25-34 characters

1. Python (Flexible, Scriptable)

Python’s line-by-line file reading keeps memory usage low—only one line is loaded at a time, perfect for large datasets.

# Extract contact numbers and save to a new file
with open('students.txt', 'r', encoding='utf-8') as infile, open('contacts_output.txt', 'w', encoding='utf-8') as outfile:
    # Optional: Write a header row for clarity
    outfile.write("StudentID,MobilePhone,HomePhone\n")
    
    for line in infile:
        # Skip empty lines to avoid errors
        if not line.strip():
            continue
        # Ensure line is long enough to contain all target fields
        if len(line) < 34:
            continue
        
        # Python uses 0-based indexing—convert your 1-based positions
        student_id = line[0:10].strip()
        mobile_phone = line[11:21].strip()  # 12th char (index 11) to 21st (index 20)
        home_phone = line[24:34].strip()    # 25th char (index24) to 34th (index33)
        
        # Write to output (use tabs instead of commas if preferred: "\t")
        outfile.write(f"{student_id},{mobile_phone},{home_phone}\n")

Why this works:

  • Minimal memory footprint: No need to load the entire 1M+ line file into RAM.
  • Easy to modify: Adjust indices if you need to extract office phone (38-47 → indices 37-47) or other fields.
  • Handles edge cases: Skips empty lines and truncated rows to prevent crashes.

2. AWK (Ultra-Fast, Command-Line)

If you prefer a no-script, terminal-based solution, AWK is ideal—it’s built for text processing and blazingly fast for large files. AWK uses 1-based indexing, which directly matches your field positions.

# Run this directly in your terminal
awk '
    length($0) >= 34 {  # Skip short/incomplete lines
        student_id = substr($0, 1, 10)
        mobile = substr($0, 12, 10)
        home = substr($0, 25, 10)
        print student_id "," mobile "," home
    }
' students.txt > contacts_output.txt

Why this works:

  • Near-native speed: AWK is written in C, so it outperforms Python for raw text processing.
  • No setup required: Just run the command—no need to write a full script.
  • Lightweight: Uses almost no additional memory, even for huge files.

Pro Tips:

  • Encoding: If your file uses a non-UTF-8 encoding (like GBK), adjust the Python encoding parameter or add -v LC_ALL=C to the AWK command for compatibility.
  • Output Format: Swap commas for \t to create a tab-separated file, or adjust the print statement to match your needs.
  • Validation: Add checks for valid phone numbers (e.g., starts with a valid country code) if needed—just extend the script/command with regex matching.

内容的提问来源于stack exchange,提问作者Dhanabalan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:19:39