如何获取CountVectorizer生成的稀疏矩阵中的词汇序列?
如何从CountVectorizer生成的稀疏矩阵反推词汇顺序?
要确定你提供的8个词汇在稀疏矩阵中的排列顺序,我们可以通过句子中出现的词汇与矩阵对应行的1的位置匹配,结合CountVectorizer的特性来推导:
步骤1:匹配已知词汇的索引
首先,我们把每个句子中包含的目标词汇和对应的矩阵行做对应:
| 句子编号 | 句子内容 | 包含的目标词汇 | 矩阵行 | 1对应的索引 | 推导的词汇 |
|---|---|---|---|---|---|
| 6 | "The intersection graph of paths in trees" | Graph | [1 0 0 0 0 0 0 0] | 0 | 索引0 → Graph |
| 8 | "Graph minors A survey" | Graph, minors | [0 1 0 0 0 0 0 0] | 1 | 索引1 → minors |
| 2、3、4 | 均包含"system" | system | 对应行的索引4为1 | 4 | 索引4 → system |
| 1、3 | 均包含"interface" | interface | 对应行的索引7为1 | 7 | 索引7 → interface |
| 1 | 包含"Human" | Human | 行中索引5为1 | 5 | 索引5 → Human |
步骤2:确定剩余词汇的索引
剩下的词汇是machine、binary、ESP,它们对应的是矩阵中全0的列(索引2、3、6)。根据CountVectorizer的默认行为,当词汇未在训练数据中出现时,它们的索引会按照**词汇的字典序(ASCII顺序)**排列:
- 索引2 → machine
- 索引3 → binary
- 索引6 → ESP
最终词汇顺序(按矩阵列索引0到7)
- 索引0: Graph
- 索引1: minors
- 索引2: machine
- 索引3: binary
- 索引4: system
- 索引5: Human
- 索引6: ESP
- 索引7: interface
验证提示
如果要100%确认顺序,最直接的方式是在代码中查看CountVectorizer的vocabulary_属性,它会返回一个{词汇: 索引}的字典:
from sklearn.feature_extraction.text import CountVectorizer vocab = ['Human', 'interface', 'machine', 'binary', 'minors', 'ESP', 'system', 'Graph'] sentences = ["Human machine interface for lab abc computer applications", "A survey of user opinion of computer system response time", "The EPS user interface management system", "Relation of user perceived response time to error measurement", "The generation of random binary unordered trees", "The intersection graph of paths in trees", "Graph minors IV Widths of trees and well quasi ordering", "Graph minors A survey"] vectorizer = CountVectorizer(vocabulary=vocab) X = vectorizer.fit_transform(sentences) print(vectorizer.vocabulary_)
运行这段代码就能得到准确的词汇-索引映射,完全不需要手动推导。
内容的提问来源于stack exchange,提问作者Eason
相关产品推荐
相关产品推荐

