使用Python对比两个文本文件,提取首列匹配的行
Python: Extract Rows from First File Where First Column Matches Entries in Second File
Hey there, let's tackle this problem. The goal is to grab all rows from File 1 where the first column (VarID) shows up in File 2. Using basic Python file operations (no external libraries required) is perfect here—fast and easy to follow.
Implementation Approach
Here's the straightforward plan:
- First, load VarIDs from File 2 into a set: Sets give us super fast lookup times, which is way more efficient than checking a list every time, especially if your files are large.
- Then process File 1 line by line:
- Keep the header row to preserve the column structure.
- For each subsequent row, split it to get the first value (VarID) and check if it’s in our set from File 2. If it is, we keep the row.
- Finally, save or print the matching rows: We can write the results to a new file or output them directly to the console.
Python Code
# Replace these paths with your actual file locations file1_path = "file1.txt" file2_path = "file2.txt" output_path = "matching_rows.txt" # Step 1: Load all valid VarIDs from File 2 into a set var_id_set = set() with open(file2_path, 'r') as file2: # Skip the header line ("VarID") next(file2) for line in file2: var_id = line.strip() if var_id: # Skip any empty lines in the file var_id_set.add(var_id) # Step 2: Collect matching rows from File 1 matching_rows = [] with open(file1_path, 'r') as file1: # Keep the header row header = next(file1).strip() matching_rows.append(header) for line in file1: cleaned_line = line.strip() if not cleaned_line: continue # Split the row by whitespace (works for spaces or tabs) row_parts = cleaned_line.split() first_column = row_parts[0] if first_column in var_id_set: matching_rows.append(cleaned_line) # Step 3: Write results to output file with open(output_path, 'w') as output_file: output_file.write('\n'.join(matching_rows)) # Optional: Print results to console for quick checking print("Matching rows found:") print('\n'.join(matching_rows))
Expected Output
Running this code with your sample files will produce the following result (saved to matching_rows.txt and printed to the console):
VarID GeneID TaxName PfamName 3810359 1327 Isochrysidaceae Methyltransf_21&Methyltransf_22 6557609 5442 Peridiniales NULL 4723299 7370 Prorocentrum PEPCK_ATP
Note: The VarIDs 5893435 and 4852156 from File 2 don’t exist in File 1, so they don’t generate any matching rows.
内容的提问来源于stack exchange,提问作者Erika
相关产品推荐
相关产品推荐

