PySpark处理《白鲸记》文本:词计数、唯一词统计、高频词及特定词计数问题求助
Hey there! Let's work through your Moby Dick text processing headaches step by step. The core issues you're hitting stem from inconsistent text cleaning, case sensitivity, and incomplete punctuation/affix removal—let's fix each task properly.
First: Nail Down the Critical Preprocessing Step
All your stats depend on clean, consistent text. Let's create a robust preprocessing pipeline first to handle case, punctuation, and suffixes like 's:
import re def clean_line(line): # Strip whitespace and skip empty lines upfront cleaned = line.strip().lower() if not cleaned: return "" # Remove possessive suffixes (e.g., "whale's" → "whale") cleaned = re.sub(r"'s\b", "", cleaned) # Remove remaining apostrophes (e.g., "don't" → "dont") cleaned = re.sub(r"'", "", cleaned) # Remove all non-alphabet characters (punctuation, numbers, symbols) cleaned = re.sub(r"[^a-zA-Z\s]", "", cleaned) return cleaned # Apply preprocessing to your raw RDD cleaned_rdd = rawMD.filter(lambda line: line.strip() != "") \ .map(clean_line) \ .flatMap(lambda line: line.split()) \ .filter(lambda word: word != "")
This pipeline ensures:
- All text is lowercase (so
Whale,WHALE, andwhaleare treated as the same word) - Possessives like
whale'sget stripped down towhale - All punctuation/symbols are removed
- No empty strings or whitespace-only entries slip through
1. Total Word Count
Now getting an accurate total is straightforward—just count the cleaned RDD entries:
total_words = cleaned_rdd.count() print(f"Total words: {total_words}")
Your original count was inflated by unprocessed punctuation and empty strings; this gives you a true count of actual words.
2. Unique Word Count
With consistent cleaning, distinct words will align much closer to Wikipedia's ~16k figure:
unique_words = cleaned_rdd.distinct().count() print(f"Unique words: {unique_words}")
The earlier 33k+ count included duplicates from case differences (Whale vs whale) and punctuation (whale vs whale,). This fix eliminates those duplicates.
3. Top 10 Most Frequent Words
You can use either RDD operations or DataFrames for this—here's both options:
RDD Approach
top_10_rdd = cleaned_rdd.map(lambda word: (word, 1)) \ .reduceByKey(lambda a, b: a + b) \ .sortBy(lambda x: x[1], ascending=False) \ .take(10) print("Top 10 frequent words:") for word, count in top_10_rdd: print(f"- {word}: {count}")
DataFrame Approach
from pyspark.sql import SparkSession from pyspark.sql.functions import desc spark = SparkSession.builder.getOrCreate() word_df = cleaned_rdd.toDF("word") top_10_df = word_df.groupBy("word") \ .count() \ .orderBy(desc("count")) \ .limit(10) top_10_df.show()
This will return clean, case-consistent results (no leading spaces like the from your original output).
4. Count Occurrences of "whale"
Since we normalized all text to lowercase and cleaned punctuation, we can match exactly on "whale" instead of partial matches:
whale_count = cleaned_rdd.filter(lambda word: word == "whale").count() print(f"'whale' appears {whale_count} times")
Your original count used "whale" in x, which would match words like whales or whalelike (skewing the count). Exact matching ensures you only count the standalone word whale, aligning with Wikipedia's ~1685 figure.
内容的提问来源于stack exchange,提问作者Demonxus

