Python中处理含字符串、数字与NaN的列统计方法问询
Got it, let's walk through how to handle this mixed-type column in pandas. The default describe() breaks because of the string value "No contact" mixed in with numbers and NaNs, and you can't replace that string with NaN or 0—totally get why that's a requirement. Here are practical approaches to generate a meaningful statistical summary:
1. Split Values by Type for Targeted Stats
The simplest way is to separate the numeric values, the "No contact" entries, and NaNs, then calculate stats for each group individually. This keeps each segment's meaning intact.
import pandas as pd import numpy as np # Example DataFrame matching your scenario df = pd.DataFrame({ 'mixed_col': [np.nan, 0, 1.32, 2, 3.4, 5, "No contact", np.nan, "No contact"] }) # Get stats for numeric values (excluding strings and NaNs) numeric_segment = df['mixed_col'].apply(lambda x: isinstance(x, (int, float))) numeric_stats = df.loc[numeric_segment, 'mixed_col'].describe() # Count "No contact" entries no_contact_count = df['mixed_col'].eq("No contact").sum() # Count NaN values nan_count = df['mixed_col'].isna().sum() # Combine all stats into a single summary full_summary = pd.concat([ numeric_stats, pd.Series({ 'no_contact_count': no_contact_count, 'nan_count': nan_count, 'no_contact_percent': (no_contact_count / len(df)) * 100, 'nan_percent': (nan_count / len(df)) * 100 }) ]) print(full_summary)
This will give you standard numeric stats (mean, min, max, etc.) plus clear counts/percentages for the non-numeric values that matter to your use case.
2. Build a Custom Summary Function
If you need to reuse this logic across multiple columns or datasets, wrap the logic into a reusable function. This lets you add extra context like unique string values or percentage breakdowns easily.
def generate_mixed_summary(series): # Isolate different value types numeric_vals = series[series.apply(lambda x: isinstance(x, (int, float)))] string_vals = series[series.apply(lambda x: isinstance(x, str))] nan_count = series.isna().sum() total_entries = len(series) # Compile stats dictionary summary_stats = {} # Add numeric stats if there are numeric values if not numeric_vals.empty: summary_stats.update(numeric_vals.describe().to_dict()) # Add string value stats summary_stats['unique_string_values'] = string_vals.unique().tolist() summary_stats['string_value_counts'] = string_vals.value_counts().to_dict() # Add NaN stats summary_stats['nan_count'] = nan_count summary_stats['nan_percentage'] = (nan_count / total_entries) * 100 if total_entries > 0 else 0 return pd.Series(summary_stats) # Apply to your column column_summary = generate_mixed_summary(df['mixed_col']) print(column_summary) # Apply to entire DataFrame (if multiple mixed columns exist) df_summary = df.apply(generate_mixed_summary) print(df_summary)
This function is flexible—if your column ever has other string values later, it'll automatically include their counts without extra code.
3. Visualize the Summary (Optional)
For better clarity, you can pair the stats with simple visualizations to highlight distributions:
- Plot a histogram for the numeric values to show their spread
- Create a bar chart for the counts of "No contact" and NaNs to show their relative frequency
Quick Example with Matplotlib:
import matplotlib.pyplot as plt # Plot numeric distribution numeric_vals = df['mixed_col'].apply(lambda x: isinstance(x, (int, float))) df.loc[numeric_vals, 'mixed_col'].dropna().hist(bins=5) plt.title('Distribution of Numeric Values') plt.show() # Plot non-numeric counts counts = pd.Series({ 'Numeric': numeric_vals.sum(), 'No contact': df['mixed_col'].eq("No contact").sum(), 'NaN': df['mixed_col'].isna().sum() }) counts.plot(kind='bar') plt.title('Breakdown of Value Types') plt.show()
Key Recommendations
- Preserve Value Meaning: Never force-cast the column to numeric (it'll turn "No contact" into NaN, which you can't do). Splitting by type is the safest way to keep each value's intent intact.
- Reuse Logic: If you work with these columns often, turn the summary function into a utility module—saves time and keeps consistency.
- Document Context: Add notes to your summary about what each segment means (e.g., "No contact = customer didn't respond") so anyone reading the stats understands the context.
内容的提问来源于stack exchange,提问作者Chris

