You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras中oov_token=True的工作原理及默认OOV词填充机制探究

How Keras Handles OOV Tokens When oov_token=True (No Explicit Token Specified)

Great question! Let's break down exactly how Keras handles out-of-vocabulary (OOV) words when you set oov_token=True but don't provide a custom token string. Even though the official docs don't spell out this detail explicitly, the behavior is consistent and easy to verify with a quick test.

1. The Default OOV Token String

When you set oov_token=True without passing a specific string, Keras automatically uses the string "<OOV>" as the default OOV marker. You can confirm this directly by checking the oov_token attribute of your Tokenizer instance:

from tensorflow.keras.preprocessing.text import Tokenizer

tokenizer = Tokenizer(oov_token=True)
print(tokenizer.oov_token)  # Output: '<OOV>'

2. Placement in the Word Index

When you call fit_on_texts() to build your vocabulary, the "<OOV>" token gets added to the word_index dictionary with the index 1. This is because Keras's Tokenizer starts counting vocabulary indices from 1 by default (index 0 is reserved for padding if you set padding=True later). Here's an example:

texts = ["cat dog", "bird fish"]
tokenizer.fit_on_texts(texts)

print(tokenizer.word_index)
# Output: {'<OOV>': 1, 'cat': 2, 'dog': 3, 'bird': 4, 'fish': 5}

3. Replacing OOV Words During Sequence Conversion

When you use texts_to_sequences() to convert raw text into numerical sequences, any word that isn't present in your trained vocabulary will be replaced with the index of "<OOV>" (which is 1 in the example above). For instance:

test_texts = ["cat hamster", "turtle snake"]
sequences = tokenizer.texts_to_sequences(test_texts)

print(sequences)
# Output: [[2, 1], [1, 1]]

Here, "hamster", "turtle", and "snake" are all out of the vocabulary, so they're replaced with the OOV token's index.

To sum it up: Keras uses "<OOV>" as the default marker when oov_token=True is set without a custom value, adds it to the vocabulary at index 1, and replaces all unseen words with this index during sequence conversion.

内容的提问来源于stack exchange,提问作者sammy ongaya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:59:09