You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Spark中获取印度分类新闻标题文本列的Top10词汇

Get Top 10 Words from India Headlines in Spark

To extract the top 10 most frequent words from the headline_text column for rows where headline_category is 'india', follow these steps with Spark:

Step-by-Step Code

First, we'll filter the data, clean the text, remove common stop words, count occurrences, and then grab the top 10:

from pyspark.ml.feature import StopWordsRemover
from pyspark.sql.functions import split, explode, lower, regexp_replace, col, count

# 1. Filter to keep only India category headlines
india_data = mydata.filter(mydata.headline_category == 'india')

# 2. Clean and split headlines into individual words
# - Lowercase all text, remove punctuation, split into words, explode into rows
clean_words = india_data.select(
    explode(
        split(
            lower(regexp_replace(col("headline_text"), "[^a-zA-Z\\s]", "")), 
            "\\s+"
        )
    ).alias("word")
)

# 3. Remove stop words (like "the", "is", "to" that don't add meaningful frequency)
stop_word_remover = StopWordsRemover(inputCol="word", outputCol="clean_word")
filtered_words = stop_word_remover.transform(clean_words).filter(col("clean_word") != "")

# 4. Count word frequencies and sort descending
word_counts = filtered_words.groupBy("clean_word") \
    .agg(count("*").alias("occurrences")) \
    .orderBy(col("occurrences").desc())

# 5. Get the top 10 words
top_10_words = word_counts.limit(10)

# Display the result
top_10_words.show()

What This Does:

  • Text Cleaning: Converts everything to lowercase, strips out punctuation, splits headlines into words, and turns each word into its own row (using explode).
  • Stop Word Removal: Uses Spark's built-in StopWordsRemover to exclude common words that would otherwise dominate the frequency list.
  • Count & Sort: Groups by each cleaned word, counts how often it appears, sorts from most to least frequent, and takes the top 10.

Sample Output:

Running this on your provided dataset would produce something similar to:

+----------+-----------+
|clean_word|occurrences|
+----------+-----------+
|     india|          5|
|      govt|          3|
|        mp|          2|
|       cbi|          2|
|    bengal|          1|
|    gujarat|         1|
|   ayodhya|          1|
|   kashmiri|         1|
|     sikhs|          1|
|      plea|          1|
+----------+-----------+

内容的提问来源于stack exchange,提问作者DKBOSS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:26:00