无法导入及使用NLTK的PunktSentenceTokenizer问题求助
Hey there! Let's break down and fix your NLTK issues step by step:
1. Fix the PunktSentenceTokenizer Import Error
Your initial attempt to import from nltk.corpus was off—PunktSentenceTokenizer belongs to the nltk.tokenize.punkt module, not the corpus module. You have two valid ways to import and initialize it:
- Direct class import:
from nltk.tokenize.punkt import PunktSentenceTokenizer tokenizer = PunktSentenceTokenizer() - Or via the nltk module namespace:
import nltk tokenizer = nltk.tokenize.punkt.PunktSentenceTokenizer()
2. Resolve the TypeError: expected string or bytes-like object
This error happens because the tokenizer.tokenize() method only accepts a single string or bytes object, but you're passing training_sentences = DataPrep.train_news['Statement']—which looks like a sequence of multiple text entries (e.g., a Pandas DataFrame column), not a single string.
You need to iterate over each text entry in the sequence and tokenize them individually. Here's how to do it cleanly:
import nltk from nltk.tokenize.punkt import PunktSentenceTokenizer # Initialize the tokenizer tokenizer = PunktSentenceTokenizer() # Your existing setup code tagged_sentences = nltk.corpus.treebank.tagged_sents() cutoff = int(.75 * len(tagged_sentences)) training_sentences = DataPrep.train_news['Statement'] # Iterate over each text entry to tokenize custom_sent_tokenizer = [] for text in training_sentences: # Ensure we're only processing string values (skip/handle non-strings as needed) if isinstance(text, str): custom_sent_tokenizer.extend(tokenizer.tokenize(text)) else: # Add logic here for non-string entries (e.g., skip empty values) pass tokenized = custom_sent_tokenizer
A quick note: If your training_sentences contains null values or non-string data, make sure to clean it first (like using DataPrep.train_news['Statement'].dropna() for Pandas) to avoid additional errors.
内容的提问来源于stack exchange,提问作者naransa

