使用SED/AWK脚本从超大规模学生文件中提取联系信息
Got it, let's tackle this problem head-on—processing a 1M+ line fixed-width file needs to be memory-efficient and fast, since loading the entire file into RAM could cause issues. Here are two robust approaches tailored for your use case:
Key Field Positions (Clarified)
First, let's map your fixed-width specs to code-friendly indices (note: different tools use 0 or 1-based indexing):
- Student ID: 1-10 characters
- Mobile Phone: 12-21 characters
- Home Phone: 25-34 characters
1. Python (Flexible, Scriptable)
Python’s line-by-line file reading keeps memory usage low—only one line is loaded at a time, perfect for large datasets.
# Extract contact numbers and save to a new file with open('students.txt', 'r', encoding='utf-8') as infile, open('contacts_output.txt', 'w', encoding='utf-8') as outfile: # Optional: Write a header row for clarity outfile.write("StudentID,MobilePhone,HomePhone\n") for line in infile: # Skip empty lines to avoid errors if not line.strip(): continue # Ensure line is long enough to contain all target fields if len(line) < 34: continue # Python uses 0-based indexing—convert your 1-based positions student_id = line[0:10].strip() mobile_phone = line[11:21].strip() # 12th char (index 11) to 21st (index 20) home_phone = line[24:34].strip() # 25th char (index24) to 34th (index33) # Write to output (use tabs instead of commas if preferred: "\t") outfile.write(f"{student_id},{mobile_phone},{home_phone}\n")
Why this works:
- Minimal memory footprint: No need to load the entire 1M+ line file into RAM.
- Easy to modify: Adjust indices if you need to extract office phone (38-47 → indices 37-47) or other fields.
- Handles edge cases: Skips empty lines and truncated rows to prevent crashes.
2. AWK (Ultra-Fast, Command-Line)
If you prefer a no-script, terminal-based solution, AWK is ideal—it’s built for text processing and blazingly fast for large files. AWK uses 1-based indexing, which directly matches your field positions.
# Run this directly in your terminal awk ' length($0) >= 34 { # Skip short/incomplete lines student_id = substr($0, 1, 10) mobile = substr($0, 12, 10) home = substr($0, 25, 10) print student_id "," mobile "," home } ' students.txt > contacts_output.txt
Why this works:
- Near-native speed: AWK is written in C, so it outperforms Python for raw text processing.
- No setup required: Just run the command—no need to write a full script.
- Lightweight: Uses almost no additional memory, even for huge files.
Pro Tips:
- Encoding: If your file uses a non-UTF-8 encoding (like GBK), adjust the Python
encodingparameter or add-v LC_ALL=Cto the AWK command for compatibility. - Output Format: Swap commas for
\tto create a tab-separated file, or adjust the print statement to match your needs. - Validation: Add checks for valid phone numbers (e.g., starts with a valid country code) if needed—just extend the script/command with regex matching.
内容的提问来源于stack exchange,提问作者Dhanabalan

