TensorFlow分词词汇丢失问题求助:texts_to_sequences输出异常
TensorFlow句子分词丢失词汇问题解决
问题重现
使用TensorFlow Keras的Tokenizer进行句子分词时,出现序列输出丢失大量词汇的情况:
from tensorflow.keras.preprocessing.text import Tokenizer #API for tokenization t = Tokenizer(num_words=4) #meant to catch most imp _ listofsentences=['Apples are fruits', 'An orange is a tasty fruit', 'Fruits are tasty!'] t.fit_on_texts(listofsentences) #processes words print(t.word_index) print(t.texts_to_sequences(listofsentences)) #arranges tokens, returns nested list
执行后,t.word_index输出正常:
{'are': 1, 'fruits': 2, 'tasty': 3, 'apples': 4, 'an': 5, 'orange': 6, 'is': 7, 'a': 8, 'fruit': 9}
但texts_to_sequences输出丢失大量词汇:
[[1, 2], [3], [2, 1, 3]]
预期输出应为:
[[4,1,2],[5,6,7,8,3,9],[2,1,3]]
错误原因
问题出在Tokenizer(num_words=4)的参数设置上:
num_words的作用是仅保留语料中出现频率最高的前N-1个词汇(词汇索引从1开始计数)。- 这里设置为4,意味着只保留排名前3的高频词(are:1、fruits:2、tasty:3),其他词汇都会被过滤,因此
texts_to_sequences只会输出这三个词的索引,其余词汇直接丢弃。 - 而
word_index属性会记录所有出现过的词汇及其索引,不受num_words参数限制,所以能看到完整的词汇字典。
解决方法
要保留所有词汇并得到预期输出,只需调整num_words参数:
- 将
num_words设置为大于语料总词汇数的值(这里总共有9个不同词汇,可设为10); - 或者直接省略
num_words参数(默认会保留所有词汇)。
修改后的代码:
from tensorflow.keras.preprocessing.text import Tokenizer # 方法1:设置足够大的num_words t = Tokenizer(num_words=10) # 方法2:省略num_words参数,默认保留所有词汇 # t = Tokenizer() listofsentences=['Apples are fruits', 'An orange is a tasty fruit', 'Fruits are tasty!'] t.fit_on_texts(listofsentences) print(t.word_index) print(t.texts_to_sequences(listofsentences))
执行后即可得到预期输出:
[[4,1,2],[5,6,7,8,3,9],[2,1,3]]
内容的提问来源于stack exchange,提问作者Anushka.N12
相关产品推荐
相关产品推荐

