You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效统计Pandas DataFrame指定国家位置频率并生成分组统计图?

Efficient Method to Aggregate and Plot User Location Frequencies

Here's a high-performance approach to group your location data into India, USA, and Others, then visualize the results—optimized for speed even with large datasets:

Step 1: Vectorized Categorization with NumPy & Pandas

Instead of slow row-wise apply() calls, we use vectorized operations (via np.select() and pandas str.contains()) which are orders of magnitude faster for large DataFrames.

Code Implementation

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# Assuming your DataFrame is already loaded as `df`

# Define grouping conditions (vectorized, regex-based)
conditions = [
    # Match any location containing "India" (case-insensitive)
    df['user_location'].str.contains(r'\bIndia\b', case=False, na=False),
    # Match US-related locations: country names, major cities, common states
    df['user_location'].str.contains(
        r'\b(United States|USA|US)\b|Washington, DC|Florida|California|New York|Texas',
        case=False, na=False
    )
]

# Corresponding group labels
choices = ['India', 'USA']

# Create the grouped column
df['location_group'] = np.select(conditions, choices, default='Others')

# Calculate frequency counts
location_frequencies = df['location_group'].value_counts()

Step 2: Visualize the Results

Use pandas' built-in plotting (powered by Matplotlib) for a quick, clean frequency bar chart:

# Plot the frequency distribution
location_frequencies.plot(
    kind='bar',
    figsize=(8, 5),
    color=['#ff6b6b', '#4ecdc4', '#45b7d1'],
    edgecolor='black'
)

# Add plot labels and title
plt.title('User Location Frequency (Grouped)', fontsize=14)
plt.xlabel('Location Group', fontsize=12)
plt.ylabel('Number of Users', fontsize=12)
plt.xticks(rotation=0)  # Keep labels horizontal for readability
plt.tight_layout()  # Adjust spacing
plt.show()

Key Optimizations & Notes

  • Vectorized Operations: np.select() and str.contains() operate on entire columns at once, avoiding the overhead of looping through each row. This is critical for datasets with tens of thousands of unique locations like yours.
  • Regex Flexibility: The regex patterns can be expanded to include more US states/cities or India-specific regions if needed. For example, add |Maharashtra to capture more Indian states.
  • Handling Missing Values: The na=False parameter ensures any missing locations are grouped into "Others".

For More Accurate Grouping (Optional)

If you need precise country matching (e.g., capturing all US cities/states without relying on regex), you can use the pycountry library to map locations to official country names. Here's a quick snippet:

import pycountry

def get_country(location):
    if pd.isna(location):
        return 'Others'
    # Check if location matches a country name
    for country in pycountry.countries:
        if country.name.lower() in location.lower():
            return 'India' if country.name == 'India' else 'USA' if country.name == 'United States' else 'Others'
    # Check for country codes or state names if needed
    return 'Others'

# Apply (note: this is slower than regex, but more accurate)
df['location_group'] = df['user_location'].apply(get_country)

This method is more accurate but slower than the regex approach—choose based on your priority between speed and precision.

内容的提问来源于stack exchange,提问作者raf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 22:57:35