You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame列中各词汇的出现频率?

Count Vocabulary Frequencies in a Pandas DataFrame Column

Counting term frequencies from a DataFrame column is a core task in text analysis, and there are straightforward ways to do this depending on how your data is structured. Here are the most common scenarios and solutions:

Scenario 1: Each row contains a single vocabulary term

If your column (let's call it nouns) has one noun per row, pandas' built-in value_counts() method is the fastest and simplest approach.

Example Code:

import pandas as pd

# Sample DataFrame matching your use case
df = pd.DataFrame({
    'nouns': ['Apple', 'Orange', 'Pear', 'Apple', 'Orange', 'Apple']
})

# Calculate frequency counts
frequency_counts = df['nouns'].value_counts()

# Print the result (formatted like your example)
print(frequency_counts)

Output:

Apple     3
Orange    2
Pear      1
Name: nouns, dtype: int64

To convert the result to a dictionary (e.g., {'Apple': 3, 'Orange': 2, ...}) for easier manipulation, add .to_dict():

freq_dict = df['nouns'].value_counts().to_dict()

Scenario 2: Each row contains multiple vocabulary terms

If your column has entries with multiple nouns (either space-separated strings or lists), you'll need to flatten the column first before counting.

Case A: Space-separated strings in each row

import pandas as pd

df = pd.DataFrame({
    'nouns': ['Apple Orange', 'Pear Apple', 'Orange Orange Apple']
})

# Split strings into individual words, then explode to create one row per word
flattened_terms = df['nouns'].str.split().explode()

# Count frequencies
frequency_counts = flattened_terms.value_counts()
print(frequency_counts)

Output:

Apple     3
Orange    3
Pear      1
Name: nouns, dtype: int64

Case B: Lists of nouns in each row

If your column already contains lists of nouns, skip the split step and use explode() directly:

df = pd.DataFrame({
    'nouns': [['Apple', 'Orange'], ['Pear', 'Apple'], ['Orange', 'Orange', 'Apple']]
})

flattened_terms = df['nouns'].explode()
frequency_counts = flattened_terms.value_counts()

Alternative: Using collections.Counter

For more flexibility (like adding custom preprocessing), Python's collections.Counter is a great option:

from collections import Counter
import pandas as pd

df = pd.DataFrame({
    'nouns': ['Apple', 'Orange', 'Pear', 'Apple', 'Orange', 'Apple']
})

# Convert the column to a list and count terms
word_counter = Counter(df['nouns'].tolist())

# Print sorted results (descending order of frequency)
for term, count in sorted(word_counter.items(), key=lambda x: x[1], reverse=True):
    print(f"{term} {count}")

Output:

Apple 3
Orange 2
Pear 1

Quick Tips:

  • Handle case insensitivity: Convert all terms to lowercase first with df['nouns'].str.lower().value_counts()
  • Clean punctuation: Use str.replace() to remove unwanted characters before counting—e.g., df['nouns'].str.replace('[^\w\s]', '', regex=True)

内容的提问来源于stack exchange,提问作者Sean

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 17:17:40