Python使用sent_tokenize统计文本文件最长/最短句子token数的问题
问题分析与修复方案
兄弟,我瞅了下你的代码,问题出在最后计算最长、最短句子token数的逻辑上——你现在的代码其实是在单个单词里找最长/最短的字符片段,而不是统计每个句子的token总数来对比,这就导致输出结果一直是1(比如单个字符的标点或者短单词)。
具体错误点拆解
- 最后那个
for x in words:循环里的words,其实是你上一轮循环中最后一个句子的token列表,不是所有句子的集合; wordpunct_tokenize(x)是把单个单词(比如"hello!")拆成["hello", "!"],然后你用max(tokenized_sentences, key=len)找的是这个拆分列表里最长的元素(比如"hello"长度是5,但如果是单个标点"!"长度就是1),完全不是你要的句子token总数。
修复后的代码
其实你已经做了最关键的一步:生成了wordcounts列表,里面已经存了每个句子的token数量!直接用这个列表找最大最小值就行,根本不需要额外循环。我还帮你优化了重复读文件的问题,修正后的完整代码如下:
from nltk.tokenize import sent_tokenize, word_tokenize from pathlib import Path def sent_length(): while True: try: file_to_open = Path(input("\nYOU CHOSE OPTION 1. Please, insert your file path: ")) # Opens and tokenize the sentences of the file with open(file_to_open) as f: sentences = sent_tokenize(f.read()) break except FileNotFoundError: print("\nFile not found. Better try again") except IsADirectoryError: print("\nIncorrect Directory path.Try again") print('\n\n This file contains', len(sentences), 'sentences in total') sent_number = 1 wordcounts = [] # 提前初始化列表,不用重复读文件 for sen in sentences: tokens = word_tokenize(sen) # Tokenize the sentence token_count = len(tokens) wordcounts.append(token_count) # 直接在这里收集token数,避免重复处理文件 # Displays the sentence number and the sentence length print('\n\nSentence', sent_number, 'contains', token_count, 'tokens') sent_number += 1 # Calculates mean sentence length average_wordcount = sum(wordcounts) / len(wordcounts) # 直接用wordcounts列表找最大最小值 longest_sen_len = max(wordcounts) shortest_sen_len = min(wordcounts) print('\n\n The longest sentence of this file contains', longest_sen_len, 'tokens') print('\n\n The shortest sentence of this file contains', shortest_sen_len, 'tokens') print('\n\nThe mean sentence length of this file is: ', average_wordcount)
优化说明
我把重复读取文件的逻辑删掉了,第一次读取文件并分割句子后,就可以在遍历句子的同时完成token统计和结果输出,既节省了IO开销,也让代码逻辑更连贯。
现在运行这个代码,就能得到你预期的结果啦!
内容的提问来源于stack exchange,提问作者Natália Resende
相关产品推荐
相关产品推荐

