You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中处理含字符串、数字与NaN的列统计方法问询

Handling Mixed-Type Column Stats in Pandas

Got it, let's walk through how to handle this mixed-type column in pandas. The default describe() breaks because of the string value "No contact" mixed in with numbers and NaNs, and you can't replace that string with NaN or 0—totally get why that's a requirement. Here are practical approaches to generate a meaningful statistical summary:

1. Split Values by Type for Targeted Stats

The simplest way is to separate the numeric values, the "No contact" entries, and NaNs, then calculate stats for each group individually. This keeps each segment's meaning intact.

import pandas as pd
import numpy as np

# Example DataFrame matching your scenario
df = pd.DataFrame({
    'mixed_col': [np.nan, 0, 1.32, 2, 3.4, 5, "No contact", np.nan, "No contact"]
})

# Get stats for numeric values (excluding strings and NaNs)
numeric_segment = df['mixed_col'].apply(lambda x: isinstance(x, (int, float)))
numeric_stats = df.loc[numeric_segment, 'mixed_col'].describe()

# Count "No contact" entries
no_contact_count = df['mixed_col'].eq("No contact").sum()

# Count NaN values
nan_count = df['mixed_col'].isna().sum()

# Combine all stats into a single summary
full_summary = pd.concat([
    numeric_stats,
    pd.Series({
        'no_contact_count': no_contact_count,
        'nan_count': nan_count,
        'no_contact_percent': (no_contact_count / len(df)) * 100,
        'nan_percent': (nan_count / len(df)) * 100
    })
])

print(full_summary)

This will give you standard numeric stats (mean, min, max, etc.) plus clear counts/percentages for the non-numeric values that matter to your use case.

2. Build a Custom Summary Function

If you need to reuse this logic across multiple columns or datasets, wrap the logic into a reusable function. This lets you add extra context like unique string values or percentage breakdowns easily.

def generate_mixed_summary(series):
    # Isolate different value types
    numeric_vals = series[series.apply(lambda x: isinstance(x, (int, float)))]
    string_vals = series[series.apply(lambda x: isinstance(x, str))]
    nan_count = series.isna().sum()
    total_entries = len(series)
    
    # Compile stats dictionary
    summary_stats = {}
    
    # Add numeric stats if there are numeric values
    if not numeric_vals.empty:
        summary_stats.update(numeric_vals.describe().to_dict())
    
    # Add string value stats
    summary_stats['unique_string_values'] = string_vals.unique().tolist()
    summary_stats['string_value_counts'] = string_vals.value_counts().to_dict()
    
    # Add NaN stats
    summary_stats['nan_count'] = nan_count
    summary_stats['nan_percentage'] = (nan_count / total_entries) * 100 if total_entries > 0 else 0
    
    return pd.Series(summary_stats)

# Apply to your column
column_summary = generate_mixed_summary(df['mixed_col'])
print(column_summary)

# Apply to entire DataFrame (if multiple mixed columns exist)
df_summary = df.apply(generate_mixed_summary)
print(df_summary)

This function is flexible—if your column ever has other string values later, it'll automatically include their counts without extra code.

3. Visualize the Summary (Optional)

For better clarity, you can pair the stats with simple visualizations to highlight distributions:

  • Plot a histogram for the numeric values to show their spread
  • Create a bar chart for the counts of "No contact" and NaNs to show their relative frequency

Quick Example with Matplotlib:

import matplotlib.pyplot as plt

# Plot numeric distribution
numeric_vals = df['mixed_col'].apply(lambda x: isinstance(x, (int, float)))
df.loc[numeric_vals, 'mixed_col'].dropna().hist(bins=5)
plt.title('Distribution of Numeric Values')
plt.show()

# Plot non-numeric counts
counts = pd.Series({
    'Numeric': numeric_vals.sum(),
    'No contact': df['mixed_col'].eq("No contact").sum(),
    'NaN': df['mixed_col'].isna().sum()
})
counts.plot(kind='bar')
plt.title('Breakdown of Value Types')
plt.show()

Key Recommendations

  • Preserve Value Meaning: Never force-cast the column to numeric (it'll turn "No contact" into NaN, which you can't do). Splitting by type is the safest way to keep each value's intent intact.
  • Reuse Logic: If you work with these columns often, turn the summary function into a utility module—saves time and keeps consistency.
  • Document Context: Add notes to your summary about what each segment means (e.g., "No contact = customer didn't respond") so anyone reading the stats understands the context.

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:33:05