You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何计算格式不统一的短信数据集的平均每句单词数?

这个问题太常见了!短信这种非正式文本的格式乱七八糟的,确实会给句子分割添不少麻烦。我自己处理类似数据集的时候也踩过坑,给你分享几个实用的解决方案,一步步搞定它:

核心思路:先准确分割句子,再统计平均单词数

问题的根源是没法识别五花八门的句末标识,所以第一步得把所有可能的句末符号都纳入识别范围,把短信拆成真正独立的句子,之后统计单词数就简单了。

步骤1:用正则表达式搞定句子分割

短信里的句末标识可不止普通句号,像省略号...、表情:)、感叹号!、问号?都是常见的句子结尾。我们可以用Python的re.split()写一个匹配模式,覆盖这些情况:

import re

# 示例短信
message = "Tiwary to rcb.battle between bang and kochi Dhawan for dc:) Warner to delhi. make it fast..."

# 正则模式:匹配句末符号(.、!、?、:)等,允许连续出现),且确保前面是字母(避免误分割缩写)
sentence_pattern = r'(?<=[a-zA-Z])[.!?:)]+\s*'
# 分割句子
sentences = re.split(sentence_pattern, message)
# 清理空字符串和前后空白
sentences = [s.strip() for s in sentences if s.strip()]

运行这段代码后,示例短信会被拆成4个清晰的句子:

['Tiwary to rcb', 'battle between bang and kochi Dhawan for dc', 'Warner to delhi', 'make it fast']

步骤2:统计单词数并计算平均值

拿到分割好的句子后,就可以逐个统计单词数,最后算出平均值:

total_words = 0
valid_sentences = 0

for sentence in sentences:
    # 按空格分割单词,过滤掉空字符串(避免连续空格的干扰)
    words = [word for word in sentence.split() if word]
    word_count = len(words)
    if word_count > 0:  # 跳过空句子
        total_words += word_count
        valid_sentences += 1

if valid_sentences > 0:
    avg_words = total_words / valid_sentences
    print(f"平均每句单词数:{avg_words:.2f}")
else:
    print("没有有效句子可统计")

针对示例短信,计算结果是4.00,完全符合预期。

进阶优化:处理更复杂的单词情况

如果你的短信里有带撇号的单词(比如don't)、数字或者连字符,用split()可能会不太准确。可以换成正则表达式提取单词:

# 匹配包含字母、数字、撇号的单词
words = re.findall(r"\b[\w']+\b", sentence)

这样能更精准地提取真正的单词,避免把标点符号当成单词的一部分。

完整流程(批量处理数据集)

如果是处理多条短信的数据集,只需要把上面的逻辑套个循环:

import re

def calculate_avg_words_per_sentence(messages):
    total_words = 0
    total_valid_sentences = 0

    for msg in messages:
        # 分割句子
        sentence_pattern = r'(?<=[a-zA-Z])[.!?:)]+\s*'
        sentences = re.split(sentence_pattern, msg)
        sentences = [s.strip() for s in sentences if s.strip()]
        
        # 统计单条短信的单词数和句子数
        for sentence in sentences:
            words = re.findall(r"\b[\w']+\b", sentence)
            word_count = len(words)
            if word_count > 0:
                total_words += word_count
                total_valid_sentences += 1

    if total_valid_sentences > 0:
        return total_words / total_valid_sentences
    else:
        return 0

# 示例数据集
dataset = [
    "Tiwary to rcb.battle between bang and kochi Dhawan for dc:) Warner to delhi. make it fast...",
    "Hello! How are you doing today?",
    "Meet me at 5pm don't be late :("
]

avg = calculate_avg_words_per_sentence(dataset)
print(f"数据集平均每句单词数:{avg:.2f}")

这样就能完美解决句末标识不统一的问题啦!

内容的提问来源于stack exchange,提问作者Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:05:53