Python遍历两个CSV文件,对比重复条目数量的实现问题
Solution to CSV Phone Number Validation Script
The original code has a couple of critical issues—most notably, nested loops will only process the first phone number from Credits correctly (since the Purchases reader gets exhausted after the first iteration), and it doesn’t properly track counts across all rows. Here’s a revised, efficient implementation that fixes these problems and meets your requirements:
Step-by-Step Approach
- First, we’ll use dictionaries to count occurrences of each phone number in both CSV files. This is an O(n + m) solution, which is far faster than nested loops.
- We’ll read each file once, extracting the 0th column (phone number) from every non-empty row.
- After counting, we’ll check every unique phone number present in either file: if the Credits count isn’t exactly twice the Purchases count, we write the number to
output.txt.
Completed Code
import csv from collections import defaultdict # Initialize dictionaries to track phone number counts (defaults to 0 for new keys) credit_counts = defaultdict(int) purchase_counts = defaultdict(int) # Count phone numbers in AllCredits.csv with open("AllCredits.csv", "r") as cred_file: creader = csv.reader(cred_file) for row in creader: # Skip empty rows to avoid index errors if row: # Strip whitespace to handle formatting inconsistencies (e.g., " 555-1234 " vs "555-1234") phone_num = row[0].strip() credit_counts[phone_num] += 1 # Count phone numbers in AllPurchases.csv with open("AllPurchases.csv", "r") as pur_file: preader = csv.reader(pur_file) for row in preader: if row: phone_num = row[0].strip() purchase_counts[phone_num] += 1 # Write invalid numbers to output.txt with open("output.txt", "w") as out_file: # Check numbers present in Credits for phone_num, cred_count in credit_counts.items(): pur_count = purchase_counts.get(phone_num, 0) if cred_count != 2 * pur_count: out_file.write(f"{phone_num}\n") # Check numbers present in Purchases but not Credits (0 credits can't equal 2*purchases unless purchases are 0) for phone_num, pur_count in purchase_counts.items(): if phone_num not in credit_counts and pur_count > 0: out_file.write(f"{phone_num}\n")
Key Improvements
- Uses
with open(...)syntax: This ensures files are automatically closed even if an error occurs, avoiding resource leaks. - Handles edge cases: Skips empty rows and normalizes phone numbers with
strip()to account for accidental whitespace. - Efficient counting: Dictionaries let us track counts in a single pass per file, instead of re-scanning Purchases for every Credits entry.
内容的提问来源于stack exchange,提问作者RachaelTheBlonde
相关产品推荐
相关产品推荐

