You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark处理《白鲸记》文本:词计数、唯一词统计、高频词及特定词计数问题求助

Hey there! Let's work through your Moby Dick text processing headaches step by step. The core issues you're hitting stem from inconsistent text cleaning, case sensitivity, and incomplete punctuation/affix removal—let's fix each task properly.

First: Nail Down the Critical Preprocessing Step

All your stats depend on clean, consistent text. Let's create a robust preprocessing pipeline first to handle case, punctuation, and suffixes like 's:

import re

def clean_line(line):
    # Strip whitespace and skip empty lines upfront
    cleaned = line.strip().lower()
    if not cleaned:
        return ""
    
    # Remove possessive suffixes (e.g., "whale's" → "whale")
    cleaned = re.sub(r"'s\b", "", cleaned)
    # Remove remaining apostrophes (e.g., "don't" → "dont")
    cleaned = re.sub(r"'", "", cleaned)
    # Remove all non-alphabet characters (punctuation, numbers, symbols)
    cleaned = re.sub(r"[^a-zA-Z\s]", "", cleaned)
    
    return cleaned

# Apply preprocessing to your raw RDD
cleaned_rdd = rawMD.filter(lambda line: line.strip() != "") \
                   .map(clean_line) \
                   .flatMap(lambda line: line.split()) \
                   .filter(lambda word: word != "")

This pipeline ensures:

  • All text is lowercase (so Whale, WHALE, and whale are treated as the same word)
  • Possessives like whale's get stripped down to whale
  • All punctuation/symbols are removed
  • No empty strings or whitespace-only entries slip through

1. Total Word Count

Now getting an accurate total is straightforward—just count the cleaned RDD entries:

total_words = cleaned_rdd.count()
print(f"Total words: {total_words}")

Your original count was inflated by unprocessed punctuation and empty strings; this gives you a true count of actual words.

2. Unique Word Count

With consistent cleaning, distinct words will align much closer to Wikipedia's ~16k figure:

unique_words = cleaned_rdd.distinct().count()
print(f"Unique words: {unique_words}")

The earlier 33k+ count included duplicates from case differences (Whale vs whale) and punctuation (whale vs whale,). This fix eliminates those duplicates.

3. Top 10 Most Frequent Words

You can use either RDD operations or DataFrames for this—here's both options:

RDD Approach

top_10_rdd = cleaned_rdd.map(lambda word: (word, 1)) \
                        .reduceByKey(lambda a, b: a + b) \
                        .sortBy(lambda x: x[1], ascending=False) \
                        .take(10)

print("Top 10 frequent words:")
for word, count in top_10_rdd:
    print(f"- {word}: {count}")

DataFrame Approach

from pyspark.sql import SparkSession
from pyspark.sql.functions import desc

spark = SparkSession.builder.getOrCreate()
word_df = cleaned_rdd.toDF("word")

top_10_df = word_df.groupBy("word") \
                   .count() \
                   .orderBy(desc("count")) \
                   .limit(10)

top_10_df.show()

This will return clean, case-consistent results (no leading spaces like the from your original output).

4. Count Occurrences of "whale"

Since we normalized all text to lowercase and cleaned punctuation, we can match exactly on "whale" instead of partial matches:

whale_count = cleaned_rdd.filter(lambda word: word == "whale").count()
print(f"'whale' appears {whale_count} times")

Your original count used "whale" in x, which would match words like whales or whalelike (skewing the count). Exact matching ensures you only count the standalone word whale, aligning with Wikipedia's ~1685 figure.


内容的提问来源于stack exchange,提问作者Demonxus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 01:42:32