You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python-Pandas数据处理:为IMDB电影数据计算基尼系数

Calculating Gini-based Diversity Score for Movie Cast Ethnicity Distribution

To compute a diversity score using a Gini-related metric for your IMDB dataset, the Gini-Simpson Index is a common and appropriate choice—it measures the probability that two randomly selected actors belong to different ethnic groups, with higher values indicating greater diversity. Here's how to implement this in pandas:

Step 1: Prepare Your Data

Assume your dataframe df has columns like tconst, ethnicity count columns (e.g., asian_count, british_count), and total_ethnicities (sum of all ethnicity counts).

Step 2: Define the Diversity Score Function

We'll calculate the Gini-Simpson Index using the formula:
diversity = 1 - sum( (ethnicity_count / total_ethnicities)² )

import pandas as pd

# Example dataframe (replace with your actual data)
data = {
    'tconst': ['tt001', 'tt002', 'tt003'],
    'asian_count': [3, 1, 5],
    'british_count': [3, 9, 5],
    'black_count': [2, 0, 0],
    'white_count': [2, 0, 0],
    'total_ethnicities': [10, 10, 10]
}
df = pd.DataFrame(data)

# Identify columns that contain ethnicity counts (adjust the filter as needed)
ethnicity_columns = [col for col in df.columns if col.endswith('_count')]

def calculate_diversity(row):
    total = row['total_ethnicities']
    if total == 0:
        return 0.0  # Handle edge case where no actors are listed
    # Calculate proportions for each ethnicity
    proportions = row[ethnicity_columns] / total
    # Compute sum of squared proportions
    sum_squared_proportions = (proportions ** 2).sum()
    # Return Gini-Simpson diversity score
    return 1 - sum_squared_proportions

# Add the diversity column to your dataframe
df['diversity'] = df.apply(calculate_diversity, axis=1)

print(df[['tconst', 'total_ethnicities', 'diversity']])

Output Explanation

For the example data:

  • tt001 has a diversity score of 0.74 (even distribution across 4 ethnicities)
  • tt002 has a score of 0.18 (90% of actors are British)
  • tt003 has a score of 0.5 (split evenly between two ethnicities)

Alternative: Standard Gini Coefficient

If you specifically need the standard Gini coefficient (which measures inequality in group sizes, where higher values mean less diversity), use this function instead:

def calculate_gini_inequality(row):
    counts = row[ethnicity_columns].values
    total = row['total_ethnicities']
    if total == 0:
        return 0.0
    # Sort counts in ascending order
    sorted_counts = sorted(counts)
    cum_counts = [0]
    for c in sorted_counts:
        cum_counts.append(cum_counts[-1] + c)
    gini = 0.0
    num_groups = len(sorted_counts)
    for i in range(num_groups):
        gini += (2 * (i + 1) - num_groups - 1) * sorted_counts[i]
    gini = gini / (total * num_groups)
    return gini

# Add inequality score (if needed, diversity would be 1 - this value)
df['gini_inequality'] = df.apply(calculate_gini_inequality, axis=1)
df['diversity_from_gini'] = 1 - df['gini_inequality']

This gives you the standard Gini coefficient for group size inequality. To use it as a diversity score, subtract it from 1 (since lower inequality equals higher diversity).

内容的提问来源于stack exchange,提问作者Shivam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:17:14