运行NLTK代码时触发TypeError: expected string or bytes-like object求助
解决NLTK文本预处理中的TypeError问题
嘿,我一眼就揪出问题所在了——你在循环里犯了个很容易踩的小坑:
在for i in range(len(sentence)):这个循环里,i是索引数字(比如0),但nltk.word_tokenize()需要的是字符串类型的句子内容,不是整数。你把索引值传给了分词函数,自然会触发TypeError: expected string or bytes-like object这个错误。
修正后的完整代码
只需要把words = nltk.word_tokenize(i)改成words = nltk.word_tokenize(sentence[i])就搞定了,完整代码如下:
import nltk from nltk.stem import PorterStemmer from nltk.corpus import stopwords # 第一次运行务必下载必要的NLTK数据集,否则会报错 nltk.download('punkt') nltk.download('stopwords') paragraph = ''' State-run Bharat Sanchar Nigam Ltd (BSNL) is readying to pay November salary in another two days, which will be raised from internal accruals and bank loans.''' sentence = nltk.sent_tokenize(paragraph) stemmer = PorterStemmer() for i in range(len(sentence)): # 这里要传入句子本身,而非循环索引i words = nltk.word_tokenize(sentence[i]) words = [stemmer.stem(word) for word in words if word not in set(stopwords.words('english'))] sentence[i] = ' '.join(words) # 可以打印看看处理后的结果 print(sentence)
额外优化小建议
其实你可以用更Pythonic的写法,直接遍历句子列表,不用手动处理索引,代码会更简洁易读:
processed_sentences = [] for sent in sentence: words = nltk.word_tokenize(sent) words = [stemmer.stem(word) for word in words if word not in set(stopwords.words('english'))] processed_sentences.append(' '.join(words))
内容的提问来源于stack exchange,提问作者jain
相关产品推荐
相关产品推荐

