Python-Pandas数据处理:为IMDB电影数据计算基尼系数
To compute a diversity score using a Gini-related metric for your IMDB dataset, the Gini-Simpson Index is a common and appropriate choice—it measures the probability that two randomly selected actors belong to different ethnic groups, with higher values indicating greater diversity. Here's how to implement this in pandas:
Step 1: Prepare Your Data
Assume your dataframe df has columns like tconst, ethnicity count columns (e.g., asian_count, british_count), and total_ethnicities (sum of all ethnicity counts).
Step 2: Define the Diversity Score Function
We'll calculate the Gini-Simpson Index using the formula:diversity = 1 - sum( (ethnicity_count / total_ethnicities)² )
import pandas as pd # Example dataframe (replace with your actual data) data = { 'tconst': ['tt001', 'tt002', 'tt003'], 'asian_count': [3, 1, 5], 'british_count': [3, 9, 5], 'black_count': [2, 0, 0], 'white_count': [2, 0, 0], 'total_ethnicities': [10, 10, 10] } df = pd.DataFrame(data) # Identify columns that contain ethnicity counts (adjust the filter as needed) ethnicity_columns = [col for col in df.columns if col.endswith('_count')] def calculate_diversity(row): total = row['total_ethnicities'] if total == 0: return 0.0 # Handle edge case where no actors are listed # Calculate proportions for each ethnicity proportions = row[ethnicity_columns] / total # Compute sum of squared proportions sum_squared_proportions = (proportions ** 2).sum() # Return Gini-Simpson diversity score return 1 - sum_squared_proportions # Add the diversity column to your dataframe df['diversity'] = df.apply(calculate_diversity, axis=1) print(df[['tconst', 'total_ethnicities', 'diversity']])
Output Explanation
For the example data:
tt001has a diversity score of 0.74 (even distribution across 4 ethnicities)tt002has a score of 0.18 (90% of actors are British)tt003has a score of 0.5 (split evenly between two ethnicities)
Alternative: Standard Gini Coefficient
If you specifically need the standard Gini coefficient (which measures inequality in group sizes, where higher values mean less diversity), use this function instead:
def calculate_gini_inequality(row): counts = row[ethnicity_columns].values total = row['total_ethnicities'] if total == 0: return 0.0 # Sort counts in ascending order sorted_counts = sorted(counts) cum_counts = [0] for c in sorted_counts: cum_counts.append(cum_counts[-1] + c) gini = 0.0 num_groups = len(sorted_counts) for i in range(num_groups): gini += (2 * (i + 1) - num_groups - 1) * sorted_counts[i] gini = gini / (total * num_groups) return gini # Add inequality score (if needed, diversity would be 1 - this value) df['gini_inequality'] = df.apply(calculate_gini_inequality, axis=1) df['diversity_from_gini'] = 1 - df['gini_inequality']
This gives you the standard Gini coefficient for group size inequality. To use it as a diversity score, subtract it from 1 (since lower inequality equals higher diversity).
内容的提问来源于stack exchange,提问作者Shivam

