如何基于Pandas DataFrame每行统计Top10高频词
Got it, I see you're looking to shift from calculating top words across your entire DataFrame to doing it row-by-row—great question! Your existing code works by concatenating all text first, but we need to adjust it to process each row independently.
First, a quick note: I noticed your full-table code references df1['description'], but your DataFrame has a comments column from the read_csv call. I'll assume that's a typo and use comments in the solution below.
Step-by-Step Solution
We can use Pandas' apply() method to run a custom function on every row of your comments column. Here's how to do it:
- Define a helper function that takes a single text string, processes it, and returns the top 10 most frequent words (along with their counts, if you want):
def get_top10_words(text): # Split text into lowercase words words = text.lower().split() # Count word frequencies and get top 10 top_words = pd.Series(words).value_counts().head(10) # Convert to a dictionary for easy storage in the DataFrame (optional) return top_words.to_dict()
- Apply this function to each row in your
commentscolumn, and store the result in a new column (e.g.,top10_words):
import pandas as pd # Your existing data loading code df1 = pd.read_csv('C:/temp/comments.csv', encoding='latin-1', names=['client', 'comments']) # Apply the helper function df1['top10_words'] = df1['comments'].apply(get_top10_words)
What This Does
apply()runsget_top10_wordson every entry in thecommentscolumn.- For each row, we split the text into lowercase words, count their occurrences, grab the top 10, and convert the result to a dictionary (so it's easy to read and work with in the DataFrame).
- If a row has fewer than 10 unique words, it will just return all the words present (no padding with empty values).
Alternative: Return Just the Word List (No Counts)
If you only want the list of top words (not their counts), modify the helper function like this:
def get_top10_word_list(text): words = text.lower().split() top_words = pd.Series(words).value_counts().head(10).index.tolist() return top_words
Then apply it the same way:
df1['top10_word_list'] = df1['comments'].apply(get_top10_word_list)
Example Output
After running the code, your DataFrame will have a new column where each entry is either a dictionary (word: count) or a list of top words, corresponding to the comments in that row.
内容的提问来源于stack exchange,提问作者John G Smith

