You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scala作业:基于Spark实现篮球关联词统计的技术问询

Hey there! Let's work through this problem together. I'll walk you through exactly how to grab the next word after "basketball", count how often each of those words shows up, and sort the results by frequency from highest to lowest using Scala with Spark.

Step-by-Step Solution

1. Read and Split the Input Text

First, we'll load the text file and split each line into individual words. We use flatMap to flatten all the words from multiple lines into a single RDD:

val lines = spark.textFile("basketball_words_only.txt")
// Split lines by any whitespace (handles spaces, tabs, etc.)
val words = lines.flatMap(line => line.split("\\s+"))

2. Grab the Word After "basketball"

The key trick here is using Spark RDD's sliding(2) method—it creates consecutive windows of 2 words, which lets us easily pair each word with the one that follows it. We then filter for pairs where the first word is "basketball" and extract the second word:

// Create (currentWord, nextWord) pairs for all consecutive words
val wordPairs = words.sliding(2).map(pair => (pair(0), pair(1)))

// Filter to keep only pairs where the first word is "basketball", then take the second word
val wordsAfterBasketball = wordPairs.filter(_._1 == "basketball").map(_._2)

Note: This automatically handles edge cases where "basketball" is the last word in the text—sliding(2) won't create a pair for it, so we avoid index-out-of-bounds errors.

3. Count Frequencies (Reduce Operation)

Next, we'll count how many times each target word appears. We use reduceByKey for efficient distributed counting:

// Map each target word to (word, 1), then sum counts per word
val wordCounts = wordsAfterBasketball.map(word => (word, 1)).reduceByKey(_ + _)

4. Sort by Frequency (Highest to Lowest)

Finally, we sort the results by the count value in descending order:

// Sort the count pairs by the second element (the frequency) in reverse order
val sortedCounts = wordCounts.sortBy(_._2, ascending = false)

// Print the results to verify
sortedCounts.collect().foreach(println)
Full Complete Code Example

Here's all the code put together in a runnable Spark application:

import org.apache.spark.sql.SparkSession

object BasketballNextWordCounter {
  def main(args: Array[String]): Unit = {
    // Initialize Spark session
    val spark = SparkSession.builder()
      .appName("BasketballNextWordCounter")
      .master("local[*]") // Remove this line for production clusters
      .getOrCreate()
    
    val lines = spark.textFile("basketball_words_only.txt")
    val words = lines.flatMap(line => line.split("\\s+"))
    
    // Extract words immediately following "basketball"
    val wordsAfterBasketball = words.sliding(2)
      .map(pair => (pair(0), pair(1)))
      .filter(_._1 == "basketball")
      .map(_._2)
    
    // Count and sort frequencies
    val sortedCounts = wordsAfterBasketball.map((_, 1))
      .reduceByKey(_ + _)
      .sortBy(_._2, ascending = false)
    
    // Show the final results
    sortedCounts.show()

    // Stop the Spark session
    spark.stop()
  }
}
Extra Tips
  • Case Insensitivity: If your text has mixed casing (like "Basketball" or "BASKETBALL"), add .map(_.toLowerCase) to the words RDD before creating pairs to catch all variations.
  • DataFrame Alternative: If you prefer working with DataFrames, you can use the lead window function to fetch the next word—though sliding is more straightforward for this simple use case.

内容的提问来源于stack exchange,提问作者Eee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:09:45