如何在Python中不使用Pandas,按两列去重(留最低价)并统计店铺数
Hey there! No worries at all—your question is totally clear and makes sense, especially when dealing with large files where pandas might not be the most efficient fit. Let's walk through a pure Python solution that handles both your tasks without relying on pandas, which should be more memory-efficient for big datasets.
解决方案步骤
1. 按Fruit+Shop分组获取最低价
We can read the file line by line (avoid loading the entire large file into memory at once) and use a dictionary to track the lowest price for each (Fruit, Shop) combination. This keeps memory usage tied only to the number of unique combinations, not the total file size.
# Store (Fruit, Shop) -> lowest price price_tracker = {} # Use 'with' to auto-manage file closure with open('your_fruit_file.txt', 'r') as f: # Skip the header line header = f.readline().strip() for line in f: # Clean up line and split into parts (handles extra spaces) parts = line.strip().split() if len(parts) != 3: continue # Skip malformed lines fruit, shop, price_str = parts # Convert price string to float for comparison price = float(price_str.replace('$', '')) # Create a composite key for the fruit-shop pair key = (fruit, shop) # Update the tracker if the key is new or current price is lower if key not in price_tracker or price < price_tracker[key]: price_tracker[key] = price # Print the first result print(header) for (fruit, shop), price in sorted(price_tracker.items()): # Format price back to $ with two decimal places print(f"{fruit} {shop} ${price:.2f}")
2. Count unique shops per fruit
We can extract fruit names from the keys of our price_tracker dictionary and use another dictionary to count how many unique shops each fruit has:
# Store Fruit -> number of unique shops shop_count = {} for fruit, shop in price_tracker.keys(): if fruit not in shop_count: shop_count[fruit] = 0 shop_count[fruit] += 1 # Print the count result print("\nFruit Shop Count:") for fruit, count in sorted(shop_count.items()): print(f"{fruit} {count}")
Why this works for large files
- Line-by-line reading: Doesn't load the entire file into memory, keeping memory usage extremely low even for GB-sized files.
- Fast dictionary lookups: Dictionary key lookups are O(1) operations, so processing stays quick regardless of file size.
- Flexible format handling: Manual line parsing lets you easily adjust for small formatting quirks (like extra spaces or price symbols) that might break pandas imports.
内容的提问来源于stack exchange,提问作者epigeneticist

