如何统计多列None值数量?Jupyter替换问号为NaN后缺失值统计返回0怎么办?
Issue 1: Replaced '?' with NaN but missing value count returns 0
First off, let's figure out why your missing value stats are showing 0 even after replacing '?' with NaN. Here are the most common fixes to try:
Double-check if the replacement actually took effect
Sometimes we forget to assign the modified DataFrame back to the original variable, or the '?' in your data might have hidden extra spaces (like ' ?'). Run these quick checks to verify:# Check unique values in a suspect column to see if NaN exists print(df['target_column'].unique()) # Verify if any NaNs exist anywhere in the DataFrame print(df.isna().any().any())Use the correct replacement syntax (and don't forget to save changes)
If you only rundf.replace('?', np.nan)without assigning the result back todf, your original data won't update. Plus, using regex handles any spaces around the question mark—a super common gotcha:import numpy as np # Replace all '?' (with or without spaces) with NaN, and save the changes df = df.replace(r'^\s*\?\s*$', np.nan, regex=True)Make sure you're using the right command to count NaNs
Don't mix up methods—here's the proper way to count missing values per column and in total:# Count NaNs for each individual column print(df.isna().sum()) # Count total number of NaNs across the entire DataFrame print(df.isna().sum().sum())
Issue 2: Count None values across multiple columns
Keep in mind that in pandas, None values in numeric columns get converted to NaN automatically, but they might stay as None in object-type columns. Here's how to count them specifically:
Count None values in a single column
# For one specific column none_count = df['target_column'].apply(lambda x: x is None).sum() print(f"Number of None values: {none_count}")Count None values across all columns
Useapplyto loop through every column and get a None count for each:# Get None count for every column in the DataFrame none_counts = df.apply(lambda col: col.apply(lambda x: x is None).sum()) print(none_counts)If you want the total number of None values across the entire dataset, just add
.sum()to the end:total_none = df.apply(lambda col: col.apply(lambda x: x is None).sum()).sum() print(f"Total None values in DataFrame: {total_none}")
If you're still stuck, sharing a tiny anonymized sample of your data and the exact commands you ran would help narrow things down further!
内容的提问来源于stack exchange,提问作者Jack

