You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas DataFrame每行统计Top10高频词

Row-Level Top 10 Word Frequency in Pandas DataFrame

Got it, I see you're looking to shift from calculating top words across your entire DataFrame to doing it row-by-row—great question! Your existing code works by concatenating all text first, but we need to adjust it to process each row independently.

First, a quick note: I noticed your full-table code references df1['description'], but your DataFrame has a comments column from the read_csv call. I'll assume that's a typo and use comments in the solution below.

Step-by-Step Solution

We can use Pandas' apply() method to run a custom function on every row of your comments column. Here's how to do it:

  1. Define a helper function that takes a single text string, processes it, and returns the top 10 most frequent words (along with their counts, if you want):
def get_top10_words(text):
    # Split text into lowercase words
    words = text.lower().split()
    # Count word frequencies and get top 10
    top_words = pd.Series(words).value_counts().head(10)
    # Convert to a dictionary for easy storage in the DataFrame (optional)
    return top_words.to_dict()
  1. Apply this function to each row in your comments column, and store the result in a new column (e.g., top10_words):
import pandas as pd

# Your existing data loading code
df1 = pd.read_csv('C:/temp/comments.csv', encoding='latin-1', names=['client', 'comments'])

# Apply the helper function
df1['top10_words'] = df1['comments'].apply(get_top10_words)

What This Does

  • apply() runs get_top10_words on every entry in the comments column.
  • For each row, we split the text into lowercase words, count their occurrences, grab the top 10, and convert the result to a dictionary (so it's easy to read and work with in the DataFrame).
  • If a row has fewer than 10 unique words, it will just return all the words present (no padding with empty values).

Alternative: Return Just the Word List (No Counts)

If you only want the list of top words (not their counts), modify the helper function like this:

def get_top10_word_list(text):
    words = text.lower().split()
    top_words = pd.Series(words).value_counts().head(10).index.tolist()
    return top_words

Then apply it the same way:

df1['top10_word_list'] = df1['comments'].apply(get_top10_word_list)

Example Output

After running the code, your DataFrame will have a new column where each entry is either a dictionary (word: count) or a list of top words, corresponding to the comments in that row.

内容的提问来源于stack exchange,提问作者John G Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:42:11