如何通过<some_model>.vocab.vectors.data行索引获取对应词汇/键?
Great question—brute-forcing through all vector items is definitely not the most efficient way to map a row index in nlp.vocab.vectors.data back to its corresponding vocabulary terms. SpaCy provides a built-in key2row dictionary that we can use to create a fast reverse lookup structure. Here's how to do it properly:
Step 1: Build a Reverse Mapping (Row Index → Vocab Keys)
First, create a dictionary where each key is a row index from the vectors data, and the value is a list of vocab hash keys that map to that row. This only needs to be computed once when you load the model:
import spacy # Load your spaCy model nlp = spacy.load('en_core_web_md') # Create reverse lookup: row index -> list of vocab keys row_to_keys = {} for vocab_key, row_idx in nlp.vocab.vectors.key2row.items(): if row_idx not in row_to_keys: row_to_keys[row_idx] = [] row_to_keys[row_idx].append(vocab_key)
Step 2: Look Up Terms by Row Index
Once you have this reverse mapping, you can quickly retrieve all terms associated with any row index. For your example with row 6:
target_row = 6 # Get all vocab keys linked to this row (empty list if no matches) matching_keys = row_to_keys.get(target_row, []) # Convert the hash keys to human-readable strings matching_terms = [nlp.vocab.strings[key] for key in matching_keys] print(f"Terms corresponding to row {target_row}: {matching_terms}")
Why This Is Better Than Brute-Force
- Speed: Building the reverse mapping takes O(N) time once, and every subsequent lookup is near-instant (O(1) for accessing the list, plus O(K) to convert keys to strings where K is the number of terms sharing the vector). Brute-forcing, by contrast, takes O(N) time for every single lookup.
- Scalability: This method works seamlessly even for large models with tens of thousands of vectors.
Important Notes
- It’s common for multiple terms to share the same vector (e.g., synonyms or terms that were assigned identical embeddings during model training), so you might get a list of terms instead of a single one.
- The
key2rowdictionary is spaCy’s source of truth for mapping vocab keys to vector rows, so this reverse mapping will always be accurate for your loaded model.
内容的提问来源于stack exchange,提问作者alexpghayes

