You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用sent_tokenize统计文本文件最长/最短句子token数的问题

问题分析与修复方案

兄弟,我瞅了下你的代码,问题出在最后计算最长、最短句子token数的逻辑上——你现在的代码其实是在单个单词里找最长/最短的字符片段,而不是统计每个句子的token总数来对比,这就导致输出结果一直是1(比如单个字符的标点或者短单词)。

具体错误点拆解

  1. 最后那个for x in words:循环里的words,其实是你上一轮循环中最后一个句子的token列表,不是所有句子的集合;
  2. wordpunct_tokenize(x)是把单个单词(比如"hello!")拆成["hello", "!"],然后你用max(tokenized_sentences, key=len)找的是这个拆分列表里最长的元素(比如"hello"长度是5,但如果是单个标点"!"长度就是1),完全不是你要的句子token总数。

修复后的代码

其实你已经做了最关键的一步:生成了wordcounts列表,里面已经存了每个句子的token数量!直接用这个列表找最大最小值就行,根本不需要额外循环。我还帮你优化了重复读文件的问题,修正后的完整代码如下:

from nltk.tokenize import sent_tokenize, word_tokenize
from pathlib import Path

def sent_length():
    while True:
        try:
            file_to_open = Path(input("\nYOU CHOSE OPTION 1. Please, insert your file path: "))
            # Opens and tokenize the sentences of the file
            with open(file_to_open) as f:
                sentences = sent_tokenize(f.read())
            break
        except FileNotFoundError:
            print("\nFile not found. Better try again")
        except IsADirectoryError:
            print("\nIncorrect Directory path.Try again")
    
    print('\n\n This file contains', len(sentences), 'sentences in total')
    sent_number = 1
    wordcounts = []  # 提前初始化列表,不用重复读文件
    for sen in sentences:
        tokens = word_tokenize(sen)  # Tokenize the sentence
        token_count = len(tokens)
        wordcounts.append(token_count)  # 直接在这里收集token数,避免重复处理文件
        # Displays the sentence number and the sentence length
        print('\n\nSentence', sent_number, 'contains', token_count, 'tokens')
        sent_number += 1

    # Calculates mean sentence length
    average_wordcount = sum(wordcounts) / len(wordcounts)
    
    # 直接用wordcounts列表找最大最小值
    longest_sen_len = max(wordcounts)
    shortest_sen_len = min(wordcounts)
    
    print('\n\n The longest sentence of this file contains', longest_sen_len, 'tokens')
    print('\n\n The shortest sentence of this file contains', shortest_sen_len, 'tokens')
    print('\n\nThe mean sentence length of this file is: ', average_wordcount)

优化说明

我把重复读取文件的逻辑删掉了,第一次读取文件并分割句子后,就可以在遍历句子的同时完成token统计和结果输出,既节省了IO开销,也让代码逻辑更连贯。

现在运行这个代码,就能得到你预期的结果啦!

内容的提问来源于stack exchange,提问作者Natália Resende

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:14:08