如何计算文本文件中各句子字符数?程序计数异常求助
问题:NLTK分句后无法正确统计句子字符数
我需要实现一个程序,将文本文件中的内容拆分为句子,并打印每个句子的字符数。我尝试用NLTK的sent_tokenize来分句,但统计结果完全不对。
尝试的代码
from collections import defaultdict import nltk from nltk.tokenize import word_tokenize from nltk.tokenize import sent_tokenize,wordpunct_tokenize import re import os import sys from pathlib import Path while True: try: file_to_open =Path(input("\nYOU SELECTED OPTION 8: CALCULATE SENTENCE LENGTH. Please, insert your file path: ")) with open(file_to_open,'r', encoding="utf-8") as f: words = sent_tokenize(f.read()) break except FileNotFoundError: print("\nFile not found. Better try again") except IsADirectoryError: print("\nIncorrect Directory path.Try again") print('\n\n This file contains',len(words),'sentences in total') wordcounts = [] caracter_count=0 sent_number=1 with open(file_to_open) as f: text = f.read() sentences = sent_tokenize(text) for sentence in sentences: if sentence.isspace() !=True: caracter_count = caracter_count + 1 print("Sentence", sent_number,'contains',caracter_count, 'characters') sent_number +=1 caracter_count = caracter_count + 1
当前异常输出
This file contains 4 sentences in total Sentence 1 contains 0 characters Sentence 2 contains 1 characters Sentence 3 contains 2 characters Sentence 4 contains 3 characters
期望输出示例
"SENTENCE 1 HAS 35 CHARACTERS" "SENTENCE 2 HAS 45 CHARACTERS" ...
嘿,我一眼就看到问题出在哪了——你代码里的caracter_count逻辑完全错了!你现在是在循环里每次给它加1,而不是去统计当前句子本身的字符数量。另外还有些冗余的操作,我帮你把代码修正并简化了:
修正后的完整代码
from collections import defaultdict import nltk from nltk.tokenize import sent_tokenize import sys from pathlib import Path # 第一次运行需要下载分句模型,加上这行避免报错 nltk.download('punkt') while True: try: file_to_open = Path(input("\nYOU SELECTED OPTION 8: CALCULATE SENTENCE LENGTH. Please, insert your file path: ")) with open(file_to_open, 'r', encoding="utf-8") as f: text = f.read() sentences = sent_tokenize(text) break except FileNotFoundError: print("\nFile not found. Better try again") except IsADirectoryError: print("\nIncorrect Directory path.Try again") print(f'\n\n This file contains {len(sentences)} sentences in total') sent_number = 1 for sentence in sentences: # 跳过全是空格的无效句子 if not sentence.isspace(): # 直接用len()获取当前句子的字符总数,这才是正确的统计方式! char_count = len(sentence) # 改成你期望的输出格式 print(f"SENTENCE {sent_number} HAS {char_count} CHARACTERS") sent_number += 1
主要改动点
- 删掉重复操作:原来你两次打开文件、两次调用
sent_tokenize,现在只需要读取一次文本就够了,省资源还更高效。 - 修正统计逻辑:把错误的
caracter_count +=1换成len(sentence),直接获取句子的实际字符数——这才是你真正需要的统计方式。 - 统一输出格式:调整成和你期望示例一致的大写格式,用f-string让代码更简洁易读。
- 添加模型下载:第一次运行时自动下载NLTK的
punkt分句模型,避免出现找不到模型的报错。
现在运行这个代码,就能得到你想要的正确结果啦!
内容的提问来源于stack exchange,提问作者Natalia Resende
相关产品推荐
相关产品推荐

