如何优化元组与文件匹配代码性能?分块方案是否正确?
First off, your chunking approach is logically correct. Splitting newlist into batches of 1000 and processing each chunk doesn’t alter the core filtering logic—you’re still iterating over every entry in newlist, just grouping them into smaller batches. This won’t change which entries end up in mylistf, so the chunking itself isn’t introducing any correctness issues.
While the chunking works, there are critical issues in both versions of your code that could lead to unexpected results or hide problems:
Newline Characters Breaking Dictionary Lookups
When you readmr_IN.dicintodc = list(f), each element includes the trailing newline (\n) from the file. But the strings you extract from the XML don’t have these newlines. So a check likei[1] in dcwill fail even if the word exists in the dictionary (e.g., "hello" != "hello\n").
Fix: Strip whitespace and filter empty lines when building your dictionary:with open("mr_IN.dic") as f: dc = {line.strip() for line in f if line.strip()} # Using a set for faster lookups!Overly Broad Try-Except Block
Yourtry-except: passcatches all exceptions, which means you’ll never know if there are malformed XML entries (with fewer than 2 values) causingIndexErrorwhen accessingi[1]. This hides useful debugging information.
Fix: Catch only the specific exception you expect:try: if i[1] in dc and i[0] not in dc: mylistf.append(i) except IndexError: # Optional: Log or print the problematic entry print(f"Skipping malformed entry: {i}") passInefficient Lookups (The Real Performance Bottleneck)
The original slowdown wasn’t due to lack of chunking—it was becausedcis a list. Checkingx in listis O(n) (linear time) for each check, which gets very slow as your dictionary grows. Convertingdcto a set turns those lookups into O(1) (constant time), which will give you a massive speedup—way more impactful than chunking.
Here’s a refined version that fixes all bugs and maximizes performance:
import re # Load dictionary into a set (strip whitespace, skip empty lines) with open("mr_IN.dic") as f: dict_set = {line.strip() for line in f if line.strip()} # Parse XML entries xml_entries = [] with open("DocumentList.xml") as f: for line in f: # Extract quoted values from each line values = re.findall(r'"([^"]*)"', line) xml_entries.append(values) # Filter desired entries filtered_entries = [] for entry in xml_entries: try: if entry[1] in dict_set and entry[0] not in dict_set: filtered_entries.append(entry) except IndexError: print(f"Skipping entry with missing values: {entry}")
Chunking might have made runtime feel more manageable (e.g., reducing memory spikes in edge cases), but the real bottleneck was linear lookups in the list-based dictionary. Using a set eliminates that bottleneck entirely, so chunking isn’t necessary for performance anymore—though you can still use it if you want to track progress or process entries in batches for other reasons.
内容的提问来源于stack exchange,提问作者shantanuo

