百万条记录文本文件的Python字符串匹配与指定列提取需求
Python Solution for Processing Large Tab-Separated Files
Hey Sai, I've got a Python solution that mirrors your existing awk/grep workflow but handles your 1M-line file efficiently—no need to load the entire file into memory, which is key for big datasets.
Here's the script:
TARGET = "2580552" INPUT_FILE = "/tmp/myfile.txt" with open(INPUT_FILE, 'r') as f: for line in f: # Strip trailing newlines and skip empty lines cleaned_line = line.strip() if not cleaned_line: continue # Split line by tabs, matching your awk -F'\t' setup fields = cleaned_line.split('\t') # Ensure we have enough columns before accessing, and check 5th column match # (Python uses 0-indexing, so 5th column = index 4) if len(fields) >= 46 and fields[4] == TARGET: # Extract columns 1,5,41,46 (convert to 0-index: 0,4,40,45) output_row = f"{fields[0]},{fields[4]},{fields[40]},{fields[45]}" print(output_row)
Key Details That Match Your Original Workflow:
- Memory Efficiency: We process the file line-by-line instead of loading all 1M rows at once, which keeps memory usage low even for massive files.
- Exact Column Matching: Just like your awk command, we only keep rows where the 5th column equals your target string—no accidental matches from other columns.
- Column Indexing Note: Python uses 0-based indexing, so the 5th column in your awk command translates to
fields[4]here, 41st column tofields[40], and 46th tofields[45]. - Robustness: The
len(fields) >= 46check prevents index errors if any lines have fewer columns than expected.
This script will produce the exact same comma-separated output as your original command, but with Python's flexibility if you ever need to extend the logic later.
内容的提问来源于stack exchange,提问作者sai_2121
相关产品推荐
相关产品推荐

