You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow分词词汇丢失问题求助:texts_to_sequences输出异常

TensorFlow句子分词丢失词汇问题解决

问题重现

使用TensorFlow Keras的Tokenizer进行句子分词时,出现序列输出丢失大量词汇的情况:

from tensorflow.keras.preprocessing.text import Tokenizer #API for tokenization

t = Tokenizer(num_words=4) #meant to catch most imp _
listofsentences=['Apples are fruits', 'An orange is a tasty fruit', 'Fruits are tasty!']
t.fit_on_texts(listofsentences) #processes words

print(t.word_index)
print(t.texts_to_sequences(listofsentences)) #arranges tokens, returns nested list

执行后,t.word_index输出正常:

{'are': 1, 'fruits': 2, 'tasty': 3, 'apples': 4, 'an': 5, 'orange': 6, 'is': 7, 'a': 8, 'fruit': 9}

但texts_to_sequences输出丢失大量词汇:

[[1, 2], [3], [2, 1, 3]]

预期输出应为:

[[4,1,2],[5,6,7,8,3,9],[2,1,3]]

错误原因

问题出在Tokenizer(num_words=4)的参数设置上:

  • num_words的作用是仅保留语料中出现频率最高的前N-1个词汇(词汇索引从1开始计数)。
  • 这里设置为4,意味着只保留排名前3的高频词(are:1、fruits:2、tasty:3),其他词汇都会被过滤,因此texts_to_sequences只会输出这三个词的索引,其余词汇直接丢弃。
  • 而word_index属性会记录所有出现过的词汇及其索引,不受num_words参数限制,所以能看到完整的词汇字典。

解决方法

要保留所有词汇并得到预期输出,只需调整num_words参数:

  1. 将num_words设置为大于语料总词汇数的值(这里总共有9个不同词汇,可设为10);
  2. 或者直接省略num_words参数(默认会保留所有词汇)。

修改后的代码:

from tensorflow.keras.preprocessing.text import Tokenizer

# 方法1:设置足够大的num_words
t = Tokenizer(num_words=10)
# 方法2:省略num_words参数,默认保留所有词汇
# t = Tokenizer()

listofsentences=['Apples are fruits', 'An orange is a tasty fruit', 'Fruits are tasty!']
t.fit_on_texts(listofsentences)

print(t.word_index)
print(t.texts_to_sequences(listofsentences))

执行后即可得到预期输出:

[[4,1,2],[5,6,7,8,3,9],[2,1,3]]

内容的提问来源于stack exchange,提问作者Anushka.N12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 11:00:49