You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用迭代器替代嵌套循环遍历spaCy tokens计算两两相似度?

Using Iterators Instead of Nested Loops for SpaCy Token Similarity Calculations

Great question! Ditching nested loops for iterators makes your code cleaner, more readable, and can even save memory when working with larger spaCy Doc objects. Let's walk through a few approaches that match your desired behavior—just like multiplying every element in [1,2,3] with every element (including itself) to get that flat list of results.

First, let's recap the nested loop approach you're already using, so we have a baseline to compare against:

import spacy

# Load a model with word vectors (required for similarity calculations)
nlp = spacy.load("en_core_web_md")
doc = nlp("Quick brown fox jumps over lazy dog")

# Original nested loop implementation
similarity_scores = []
for token1 in doc:
    for token2 in doc:
        similarity_scores.append(token1.similarity(token2))

Approach 1: Use itertools.product for Cartesian Product Pairs

The itertools.product function is perfect here—it generates the Cartesian product of your Doc with itself, which is exactly all possible (token1, token2) pairs (including every token paired with itself). This replaces the nested loops with a clean iterator-based approach:

from itertools import product
import spacy

nlp = spacy.load("en_core_web_md")
doc = nlp("Quick brown fox jumps over lazy dog")

# Generate all token pairs as an iterator
token_pairs = product(doc, doc)

# Calculate similarities using a list comprehension (or generator expression)
similarity_scores = [t1.similarity(t2) for t1, t2 in token_pairs]

# If you want a memory-efficient generator instead of a list (great for large docs)
similarity_generator = (t1.similarity(t2) for t1, t2 in product(doc, doc))

This works exactly like your nested loops—you'll get the same order of scores, but with more concise code. The generator version is especially handy if you don't need all scores stored in memory at once; you can iterate through it one score at a time.

Approach 2: Nested List Comprehension (Iterator Under the Hood)

If you prefer to skip importing itertools, a nested list comprehension achieves the same result with even less code. Under the hood, it still uses iterators, and it's functionally identical to your original nested loops:

import spacy

nlp = spacy.load("en_core_web_md")
doc = nlp("Quick brown fox jumps over lazy dog")

similarity_scores = [t1.similarity(t2) for t1 in doc for t2 in doc]

This is a more compact way to write the same logic, and it's easy to read once you're familiar with list comprehension syntax.

Quick Note About SpaCy Similarity

Just a reminder: To use token.similarity(), you need to load a spaCy model that includes pre-trained word vectors (like en_core_web_md or en_core_web_lg). If you use a smaller model like en_core_web_sm, you'll get an error because it doesn't include word vectors for similarity calculations.

内容的提问来源于stack exchange,提问作者Davtho1983

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:05:24