Keras Tokenizer仅对CSV文件首行执行字符级分词问题求助
Hey there! Let's figure out why your Tokenizer is only processing what looks like the first line—it's actually not the first row of your CSV, but the column name of your DataFrame, and here's how to fix it:
The Root Cause
When you pass test_dataset (a pandas DataFrame with one column) to fit_on_texts(), Keras Tokenizer iterates over the columns of the DataFrame instead of the rows in your host column. That means it's only processing the string 'host' (the column label) instead of all the actual host values in your CSV.
Fixed Code
import pandas as pd from keras.preprocessing.text import Tokenizer # Read the CSV and extract the 'host' column as a Series fields = ['host'] test_dataset = pd.read_csv('dga_data.csv', usecols=fields) host_texts = test_dataset['host'] # This is the key—grab the actual column data # Initialize Tokenizer (char_level=True handles character-level tokenization automatically) test_dataset_tok = Tokenizer(char_level=True, oov_token=True) # Fit on the actual host texts, not the DataFrame test_dataset_tok.fit_on_texts(host_texts) # Convert all host texts to sequences test_dataset_sequences = test_dataset_tok.texts_to_sequences(host_texts) # Verify your results print(test_dataset_tok) print(test_dataset_sequences) print(test_dataset_tok.word_index)
Quick Notes
host_textsis a pandas Series containing every value from yourhostcolumn—this is the iterable of strings thatfit_on_texts()expects.- Since you're using
char_level=True, thesplitparameter is ignored (Tokenizer splits every character automatically), so I removed it to avoid confusion. oov_token=Trueensures any unseen characters later get a dedicated token, which is perfect for your DGA detection use case.
内容的提问来源于stack exchange,提问作者Rholmes
相关产品推荐
相关产品推荐

