You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MultiLabelBinarizer能否实现多标签值计数?附DataFrame场景示例

Perfect question! The MultiLabelBinarizer from scikit-learn is built to output binary (0/1) flags for whether a label exists in a list, but we can easily extend it to get the frequency counts you need. Here's a straightforward way to do this, which works great even for large datasets of code lists:

Step 1: Set up your data and imports

First, let's import the necessary libraries and create your sample DataFrame:

import pandas as pd
from sklearn.preprocessing import MultiLabelBinarizer
from collections import Counter

# Sample DataFrame with list values
df = pd.DataFrame({
    'a': [['earth','mars','earth','moon'], ['jupiter','pluto','sun']]
})

Step 2: Use MultiLabelBinarizer to get all unique labels

We'll use MultiLabelBinarizer to identify every unique label across all rows. This ensures we have a consistent set of columns, even if some rows don't contain certain labels:

mlb = MultiLabelBinarizer()
mlb.fit(df['a'])
all_unique_labels = mlb.classes_  # Output: ['earth', 'jupiter', 'mars', 'moon', 'pluto', 'sun']

Step 3: Calculate label frequencies for each row

Next, we'll create a helper function to count how many times each label appears in a list, then align those counts with our full set of unique labels:

def count_label_occurrences(label_list):
    # Count frequency of each label in the list
    label_counts = Counter(label_list)
    # Convert to a Series with all unique labels, filling missing values with 0
    return pd.Series(label_counts, index=all_unique_labels).fillna(0).astype(int)

# Apply the function to every row in column 'a'
count_result_df = df['a'].apply(count_label_occurrences)

Step 4: Final result

The output will match exactly what you're looking for:

earth  jupiter  mars  moon  pluto  sun
0      2        0     1     1      0    0
1      0        1     0     0      1    1

If you want to reorder the columns to match your example's sequence, just reindex the DataFrame:

count_result_df = count_result_df[['earth', 'mars', 'moon', 'sun', 'jupiter', 'pluto']]

This approach is efficient for large datasets, as it leverages pandas' optimized operations and keeps your label set consistent using MultiLabelBinarizer.

内容的提问来源于stack exchange,提问作者Grigor Carran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:42:32