如何计算格式不统一的短信数据集的平均每句单词数?
这个问题太常见了!短信这种非正式文本的格式乱七八糟的,确实会给句子分割添不少麻烦。我自己处理类似数据集的时候也踩过坑,给你分享几个实用的解决方案,一步步搞定它:
核心思路:先准确分割句子,再统计平均单词数
问题的根源是没法识别五花八门的句末标识,所以第一步得把所有可能的句末符号都纳入识别范围,把短信拆成真正独立的句子,之后统计单词数就简单了。
步骤1:用正则表达式搞定句子分割
短信里的句末标识可不止普通句号,像省略号...、表情:)、感叹号!、问号?都是常见的句子结尾。我们可以用Python的re.split()写一个匹配模式,覆盖这些情况:
import re # 示例短信 message = "Tiwary to rcb.battle between bang and kochi Dhawan for dc:) Warner to delhi. make it fast..." # 正则模式:匹配句末符号(.、!、?、:)等,允许连续出现),且确保前面是字母(避免误分割缩写) sentence_pattern = r'(?<=[a-zA-Z])[.!?:)]+\s*' # 分割句子 sentences = re.split(sentence_pattern, message) # 清理空字符串和前后空白 sentences = [s.strip() for s in sentences if s.strip()]
运行这段代码后,示例短信会被拆成4个清晰的句子:
['Tiwary to rcb', 'battle between bang and kochi Dhawan for dc', 'Warner to delhi', 'make it fast']
步骤2:统计单词数并计算平均值
拿到分割好的句子后,就可以逐个统计单词数,最后算出平均值:
total_words = 0 valid_sentences = 0 for sentence in sentences: # 按空格分割单词,过滤掉空字符串(避免连续空格的干扰) words = [word for word in sentence.split() if word] word_count = len(words) if word_count > 0: # 跳过空句子 total_words += word_count valid_sentences += 1 if valid_sentences > 0: avg_words = total_words / valid_sentences print(f"平均每句单词数:{avg_words:.2f}") else: print("没有有效句子可统计")
针对示例短信,计算结果是4.00,完全符合预期。
进阶优化:处理更复杂的单词情况
如果你的短信里有带撇号的单词(比如don't)、数字或者连字符,用split()可能会不太准确。可以换成正则表达式提取单词:
# 匹配包含字母、数字、撇号的单词 words = re.findall(r"\b[\w']+\b", sentence)
这样能更精准地提取真正的单词,避免把标点符号当成单词的一部分。
完整流程(批量处理数据集)
如果是处理多条短信的数据集,只需要把上面的逻辑套个循环:
import re def calculate_avg_words_per_sentence(messages): total_words = 0 total_valid_sentences = 0 for msg in messages: # 分割句子 sentence_pattern = r'(?<=[a-zA-Z])[.!?:)]+\s*' sentences = re.split(sentence_pattern, msg) sentences = [s.strip() for s in sentences if s.strip()] # 统计单条短信的单词数和句子数 for sentence in sentences: words = re.findall(r"\b[\w']+\b", sentence) word_count = len(words) if word_count > 0: total_words += word_count total_valid_sentences += 1 if total_valid_sentences > 0: return total_words / total_valid_sentences else: return 0 # 示例数据集 dataset = [ "Tiwary to rcb.battle between bang and kochi Dhawan for dc:) Warner to delhi. make it fast...", "Hello! How are you doing today?", "Meet me at 5pm don't be late :(" ] avg = calculate_avg_words_per_sentence(dataset) print(f"数据集平均每句单词数:{avg:.2f}")
这样就能完美解决句末标识不统一的问题啦!
内容的提问来源于stack exchange,提问作者Student
相关产品推荐
相关产品推荐

