You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将自定义ChainingDict与词与行索引应用集成并计时?

Integrating ChainingDict into Your Concordance Program

First, we need to fix the missing __iter__ method in your ChainingDict (required to sort keys later), then replace the defaultdict with your custom dictionary, and finally add timing logic to measure performance.

Step 1: Implement the __iter__ Method in ChainingDict

Your ChainingDict needs an iterator to loop through all keys. Add this code to fill in the missing __iter__ method:

def __iter__(self):
    """Iterates over the keys of the dictionary"""
    # Traverse each linked list in the hash table
    for linked_list in self._table:
        # Assuming your UnorderedList uses a 'head' attribute to start the list
        current_node = linked_list.head
        while current_node is not None:
            # Yield the key from each Entry stored in the node
            yield current_node.data.getKey()
            current_node = current_node.next

Note: Adjust this if your UnorderedList uses different attribute names (e.g., if nodes store entries in a field other than data).

Step 2: Replace defaultdict with ChainingDict in the Concordance Function

Modify your concordance code to use ChainingDict instead of defaultdict(list). Since ChainingDict doesn’t auto-initialize lists for missing keys, we’ll handle that manually:

import re
from chaining_dict import ChainingDict  # Import your custom dictionary class

def concordance():
    wordConcordanceDict = ChainingDict()  # Use ChainingDict instead of defaultdict
    with open('stop_words_small.txt') as sw:
        stop_words = set(line.strip() for line in sw)
    with open('small_file.txt') as f:
        for line_number, line in enumerate(f, 1):
            words = (re.sub(r'[^\w\s]','',word).lower() for word in line.split())
            good_words = (word for word in words if word not in stop_words)
            for word in good_words:
                # Get existing line numbers or create a new list if the word is new
                current_lines = wordConcordanceDict[word]
                if not current_lines:
                    current_lines = []
                current_lines.append(line_number)
                # Update the dictionary with the modified list
                wordConcordanceDict[word] = current_lines
    # Sort and print the final concordance results
    for word in sorted(wordConcordanceDict):
        print('{}: {}'.format(word, ' '.join(map(str, wordConcordanceDict[word]))))

Step 3: Add Timing to Your Program

To measure how long the program takes to run, use Python’s time.perf_counter() (it’s more precise than time.time() for short durations). Wrap your concordance() call with timing logic:

import time

if __name__ == "__main__":
    start_time = time.perf_counter()
    concordance()
    end_time = time.perf_counter()
    print(f"\nTotal execution time: {end_time - start_time:.6f} seconds")

Comparing Performance (Optional)

If you want to compare your ChainingDict with the original defaultdict version, create separate functions for each and time both:

import re
from collections import defaultdict
from chaining_dict import ChainingDict
import time

def concordance_defaultdict():
    wordConcordanceDict = defaultdict(list)
    with open('stop_words_small.txt') as sw:
        stop_words = set(line.strip() for line in sw)
    with open('small_file.txt') as f:
        for line_num, line in enumerate(f, 1):
            words = (re.sub(r'[^\w\s]','',w.lower()) for w in line.split())
            good_words = (w for w in words if w not in stop_words)
            for word in good_words:
                wordConcordanceDict[word].append(line_num)
    for word in sorted(wordConcordanceDict):
        print(f"{word}: {' '.join(map(str, wordConcordanceDict[word]))}")

def concordance_chainingdict():
    # Reuse the modified function from Step 2 here
    wordConcordanceDict = ChainingDict()
    with open('stop_words_small.txt') as sw:
        stop_words = set(line.strip() for line in sw)
    with open('small_file.txt') as f:
        for line_num, line in enumerate(f, 1):
            words = (re.sub(r'[^\w\s]','',w.lower()) for w in line.split())
            good_words = (w for w in words if w not in stop_words)
            for word in good_words:
                current_lines = wordConcordanceDict[word]
                if not current_lines:
                    current_lines = []
                current_lines.append(line_num)
                wordConcordanceDict[word] = current_lines
    for word in sorted(wordConcordanceDict):
        print(f"{word}: {' '.join(map(str, wordConcordanceDict[word]))}")

if __name__ == "__main__":
    # Time defaultdict version
    print("Running defaultdict version...")
    start = time.perf_counter()
    concordance_defaultdict()
    end = time.perf_counter()
    print(f"Defaultdict time: {end - start:.6f}s\n")

    # Time ChainingDict version
    print("Running ChainingDict version...")
    start = time.perf_counter()
    concordance_chainingdict()
    end = time.perf_counter()
    print(f"ChainingDict time: {end - start:.6f}s")

Key Notes

  • Ensure your Entry class has getKey() and getValue() methods (your existing ChainingDict code relies on these, so they should already exist).
  • If your UnorderedList has an __iter__ method, you can simplify the ChainingDict.__iter__ to loop directly over the linked list instead of traversing via head.

内容的提问来源于stack exchange,提问作者Fatima Tabasum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:32:47