如何将自定义ChainingDict与词与行索引应用集成并计时?
First, we need to fix the missing __iter__ method in your ChainingDict (required to sort keys later), then replace the defaultdict with your custom dictionary, and finally add timing logic to measure performance.
Step 1: Implement the __iter__ Method in ChainingDict
Your ChainingDict needs an iterator to loop through all keys. Add this code to fill in the missing __iter__ method:
def __iter__(self): """Iterates over the keys of the dictionary""" # Traverse each linked list in the hash table for linked_list in self._table: # Assuming your UnorderedList uses a 'head' attribute to start the list current_node = linked_list.head while current_node is not None: # Yield the key from each Entry stored in the node yield current_node.data.getKey() current_node = current_node.next
Note: Adjust this if your UnorderedList uses different attribute names (e.g., if nodes store entries in a field other than data).
Step 2: Replace defaultdict with ChainingDict in the Concordance Function
Modify your concordance code to use ChainingDict instead of defaultdict(list). Since ChainingDict doesn’t auto-initialize lists for missing keys, we’ll handle that manually:
import re from chaining_dict import ChainingDict # Import your custom dictionary class def concordance(): wordConcordanceDict = ChainingDict() # Use ChainingDict instead of defaultdict with open('stop_words_small.txt') as sw: stop_words = set(line.strip() for line in sw) with open('small_file.txt') as f: for line_number, line in enumerate(f, 1): words = (re.sub(r'[^\w\s]','',word).lower() for word in line.split()) good_words = (word for word in words if word not in stop_words) for word in good_words: # Get existing line numbers or create a new list if the word is new current_lines = wordConcordanceDict[word] if not current_lines: current_lines = [] current_lines.append(line_number) # Update the dictionary with the modified list wordConcordanceDict[word] = current_lines # Sort and print the final concordance results for word in sorted(wordConcordanceDict): print('{}: {}'.format(word, ' '.join(map(str, wordConcordanceDict[word]))))
Step 3: Add Timing to Your Program
To measure how long the program takes to run, use Python’s time.perf_counter() (it’s more precise than time.time() for short durations). Wrap your concordance() call with timing logic:
import time if __name__ == "__main__": start_time = time.perf_counter() concordance() end_time = time.perf_counter() print(f"\nTotal execution time: {end_time - start_time:.6f} seconds")
Comparing Performance (Optional)
If you want to compare your ChainingDict with the original defaultdict version, create separate functions for each and time both:
import re from collections import defaultdict from chaining_dict import ChainingDict import time def concordance_defaultdict(): wordConcordanceDict = defaultdict(list) with open('stop_words_small.txt') as sw: stop_words = set(line.strip() for line in sw) with open('small_file.txt') as f: for line_num, line in enumerate(f, 1): words = (re.sub(r'[^\w\s]','',w.lower()) for w in line.split()) good_words = (w for w in words if w not in stop_words) for word in good_words: wordConcordanceDict[word].append(line_num) for word in sorted(wordConcordanceDict): print(f"{word}: {' '.join(map(str, wordConcordanceDict[word]))}") def concordance_chainingdict(): # Reuse the modified function from Step 2 here wordConcordanceDict = ChainingDict() with open('stop_words_small.txt') as sw: stop_words = set(line.strip() for line in sw) with open('small_file.txt') as f: for line_num, line in enumerate(f, 1): words = (re.sub(r'[^\w\s]','',w.lower()) for w in line.split()) good_words = (w for w in words if w not in stop_words) for word in good_words: current_lines = wordConcordanceDict[word] if not current_lines: current_lines = [] current_lines.append(line_num) wordConcordanceDict[word] = current_lines for word in sorted(wordConcordanceDict): print(f"{word}: {' '.join(map(str, wordConcordanceDict[word]))}") if __name__ == "__main__": # Time defaultdict version print("Running defaultdict version...") start = time.perf_counter() concordance_defaultdict() end = time.perf_counter() print(f"Defaultdict time: {end - start:.6f}s\n") # Time ChainingDict version print("Running ChainingDict version...") start = time.perf_counter() concordance_chainingdict() end = time.perf_counter() print(f"ChainingDict time: {end - start:.6f}s")
Key Notes
- Ensure your
Entryclass hasgetKey()andgetValue()methods (your existingChainingDictcode relies on these, so they should already exist). - If your
UnorderedListhas an__iter__method, you can simplify theChainingDict.__iter__to loop directly over the linked list instead of traversing viahead.
内容的提问来源于stack exchange,提问作者Fatima Tabasum

