如何统计DataFrame列中各词汇的出现频率?
Counting term frequencies from a DataFrame column is a core task in text analysis, and there are straightforward ways to do this depending on how your data is structured. Here are the most common scenarios and solutions:
Scenario 1: Each row contains a single vocabulary term
If your column (let's call it nouns) has one noun per row, pandas' built-in value_counts() method is the fastest and simplest approach.
Example Code:
import pandas as pd # Sample DataFrame matching your use case df = pd.DataFrame({ 'nouns': ['Apple', 'Orange', 'Pear', 'Apple', 'Orange', 'Apple'] }) # Calculate frequency counts frequency_counts = df['nouns'].value_counts() # Print the result (formatted like your example) print(frequency_counts)
Output:
Apple 3 Orange 2 Pear 1 Name: nouns, dtype: int64
To convert the result to a dictionary (e.g., {'Apple': 3, 'Orange': 2, ...}) for easier manipulation, add .to_dict():
freq_dict = df['nouns'].value_counts().to_dict()
Scenario 2: Each row contains multiple vocabulary terms
If your column has entries with multiple nouns (either space-separated strings or lists), you'll need to flatten the column first before counting.
Case A: Space-separated strings in each row
import pandas as pd df = pd.DataFrame({ 'nouns': ['Apple Orange', 'Pear Apple', 'Orange Orange Apple'] }) # Split strings into individual words, then explode to create one row per word flattened_terms = df['nouns'].str.split().explode() # Count frequencies frequency_counts = flattened_terms.value_counts() print(frequency_counts)
Output:
Apple 3 Orange 3 Pear 1 Name: nouns, dtype: int64
Case B: Lists of nouns in each row
If your column already contains lists of nouns, skip the split step and use explode() directly:
df = pd.DataFrame({ 'nouns': [['Apple', 'Orange'], ['Pear', 'Apple'], ['Orange', 'Orange', 'Apple']] }) flattened_terms = df['nouns'].explode() frequency_counts = flattened_terms.value_counts()
Alternative: Using collections.Counter
For more flexibility (like adding custom preprocessing), Python's collections.Counter is a great option:
from collections import Counter import pandas as pd df = pd.DataFrame({ 'nouns': ['Apple', 'Orange', 'Pear', 'Apple', 'Orange', 'Apple'] }) # Convert the column to a list and count terms word_counter = Counter(df['nouns'].tolist()) # Print sorted results (descending order of frequency) for term, count in sorted(word_counter.items(), key=lambda x: x[1], reverse=True): print(f"{term} {count}")
Output:
Apple 3 Orange 2 Pear 1
Quick Tips:
- Handle case insensitivity: Convert all terms to lowercase first with
df['nouns'].str.lower().value_counts() - Clean punctuation: Use
str.replace()to remove unwanted characters before counting—e.g.,df['nouns'].str.replace('[^\w\s]', '', regex=True)
内容的提问来源于stack exchange,提问作者Sean

