如何高效统计Pandas DataFrame指定国家位置频率并生成分组统计图?
Here's a high-performance approach to group your location data into India, USA, and Others, then visualize the results—optimized for speed even with large datasets:
Step 1: Vectorized Categorization with NumPy & Pandas
Instead of slow row-wise apply() calls, we use vectorized operations (via np.select() and pandas str.contains()) which are orders of magnitude faster for large DataFrames.
Code Implementation
import pandas as pd import numpy as np import matplotlib.pyplot as plt # Assuming your DataFrame is already loaded as `df` # Define grouping conditions (vectorized, regex-based) conditions = [ # Match any location containing "India" (case-insensitive) df['user_location'].str.contains(r'\bIndia\b', case=False, na=False), # Match US-related locations: country names, major cities, common states df['user_location'].str.contains( r'\b(United States|USA|US)\b|Washington, DC|Florida|California|New York|Texas', case=False, na=False ) ] # Corresponding group labels choices = ['India', 'USA'] # Create the grouped column df['location_group'] = np.select(conditions, choices, default='Others') # Calculate frequency counts location_frequencies = df['location_group'].value_counts()
Step 2: Visualize the Results
Use pandas' built-in plotting (powered by Matplotlib) for a quick, clean frequency bar chart:
# Plot the frequency distribution location_frequencies.plot( kind='bar', figsize=(8, 5), color=['#ff6b6b', '#4ecdc4', '#45b7d1'], edgecolor='black' ) # Add plot labels and title plt.title('User Location Frequency (Grouped)', fontsize=14) plt.xlabel('Location Group', fontsize=12) plt.ylabel('Number of Users', fontsize=12) plt.xticks(rotation=0) # Keep labels horizontal for readability plt.tight_layout() # Adjust spacing plt.show()
Key Optimizations & Notes
- Vectorized Operations:
np.select()andstr.contains()operate on entire columns at once, avoiding the overhead of looping through each row. This is critical for datasets with tens of thousands of unique locations like yours. - Regex Flexibility: The regex patterns can be expanded to include more US states/cities or India-specific regions if needed. For example, add
|Maharashtrato capture more Indian states. - Handling Missing Values: The
na=Falseparameter ensures any missing locations are grouped into "Others".
For More Accurate Grouping (Optional)
If you need precise country matching (e.g., capturing all US cities/states without relying on regex), you can use the pycountry library to map locations to official country names. Here's a quick snippet:
import pycountry def get_country(location): if pd.isna(location): return 'Others' # Check if location matches a country name for country in pycountry.countries: if country.name.lower() in location.lower(): return 'India' if country.name == 'India' else 'USA' if country.name == 'United States' else 'Others' # Check for country codes or state names if needed return 'Others' # Apply (note: this is slower than regex, but more accurate) df['location_group'] = df['user_location'].apply(get_country)
This method is more accurate but slower than the regex approach—choose based on your priority between speed and precision.
内容的提问来源于stack exchange,提问作者raf

