如何用迭代器替代嵌套循环遍历spaCy tokens计算两两相似度?
Great question! Ditching nested loops for iterators makes your code cleaner, more readable, and can even save memory when working with larger spaCy Doc objects. Let's walk through a few approaches that match your desired behavior—just like multiplying every element in [1,2,3] with every element (including itself) to get that flat list of results.
First, let's recap the nested loop approach you're already using, so we have a baseline to compare against:
import spacy # Load a model with word vectors (required for similarity calculations) nlp = spacy.load("en_core_web_md") doc = nlp("Quick brown fox jumps over lazy dog") # Original nested loop implementation similarity_scores = [] for token1 in doc: for token2 in doc: similarity_scores.append(token1.similarity(token2))
Approach 1: Use itertools.product for Cartesian Product Pairs
The itertools.product function is perfect here—it generates the Cartesian product of your Doc with itself, which is exactly all possible (token1, token2) pairs (including every token paired with itself). This replaces the nested loops with a clean iterator-based approach:
from itertools import product import spacy nlp = spacy.load("en_core_web_md") doc = nlp("Quick brown fox jumps over lazy dog") # Generate all token pairs as an iterator token_pairs = product(doc, doc) # Calculate similarities using a list comprehension (or generator expression) similarity_scores = [t1.similarity(t2) for t1, t2 in token_pairs] # If you want a memory-efficient generator instead of a list (great for large docs) similarity_generator = (t1.similarity(t2) for t1, t2 in product(doc, doc))
This works exactly like your nested loops—you'll get the same order of scores, but with more concise code. The generator version is especially handy if you don't need all scores stored in memory at once; you can iterate through it one score at a time.
Approach 2: Nested List Comprehension (Iterator Under the Hood)
If you prefer to skip importing itertools, a nested list comprehension achieves the same result with even less code. Under the hood, it still uses iterators, and it's functionally identical to your original nested loops:
import spacy nlp = spacy.load("en_core_web_md") doc = nlp("Quick brown fox jumps over lazy dog") similarity_scores = [t1.similarity(t2) for t1 in doc for t2 in doc]
This is a more compact way to write the same logic, and it's easy to read once you're familiar with list comprehension syntax.
Quick Note About SpaCy Similarity
Just a reminder: To use token.similarity(), you need to load a spaCy model that includes pre-trained word vectors (like en_core_web_md or en_core_web_lg). If you use a smaller model like en_core_web_sm, you'll get an error because it doesn't include word vectors for similarity calculations.
内容的提问来源于stack exchange,提问作者Davtho1983

