pd.to_numeric转换报错ValueError:无法解析‘notpresent’求解决方案
Fixing Numeric Conversion Error in Your Python Data Processing Code
Hey there, let's break down what's going wrong and fix that ValueError you're seeing.
What's Causing the Error?
The problem is straightforward: your bgr column contains non-numeric strings like "notpresent" (complete with extra double quotes!), and pd.to_numeric() can't parse these into floating-point numbers by default. That's exactly why you're getting the Unable to parse string error.
Step-by-Step Fix
Let's walk through the fixes one by one:
- Strip Unwanted Double Quotes
First, we need to remove those extra double quotes from all string columns—they're blocking proper value conversion:
# Remove double quotes from any object-type columns df = df.apply(lambda col: col.str.replace('"', '') if col.dtype == 'object' else col)
- Convert Columns Gracefully
Use theerrors='coerce'argument withpd.to_numeric()—this will turn any unparseable strings intoNaNinstead of crashing your code:
# Convert columns to float, turning bad values into NaN df['bgr'] = pd.to_numeric(df['bgr'], errors='coerce', downcast='float') df['wc'] = pd.to_numeric(df['wc'], errors='coerce', downcast='float') df['rc'] = pd.to_numeric(df['rc'], errors='coerce', downcast='float')
- Handle Missing Values (Optional but Important)
Now you'll haveNaNvalues in your dataframe. You can either drop rows with missing values or fill them with a statistic like the mean/median (filling is usually better for retaining data):
# Option 1: Drop rows with any NaN values df.dropna(axis=0, how='any', inplace=True) # Option 2: Fill NaNs with column means # df.fillna(df.mean(numeric_only=True), inplace=True)
Full Fixed Code
Here's your complete code with all the fixes applied:
import math import pandas as pd import matplotlib.pyplot as plt import matplotlib from sklearn import preprocessing # Load the dataset df = pd.read_html('https://github.com/authman/DAT210x/blob/master/Module4/Datasets/kidney_disease.csv')[0] # Drop the id column df.drop('id', axis=1, inplace=True) # Keep only the specified columns df = df[['bgr', 'wc', 'rc']] # Clean up double quotes from string columns df = df.apply(lambda col: col.str.replace('"', '') if col.dtype == 'object' else col) # Convert columns to numeric, coercing bad values to NaN df['bgr'] = pd.to_numeric(df['bgr'], errors='coerce', downcast='float') df['wc'] = pd.to_numeric(df['wc'], errors='coerce', downcast='float') df['rc'] = pd.to_numeric(df['rc'], errors='coerce', downcast='float') # Handle missing values (choose one option) df.dropna(axis=0, how='any', inplace=True) # df.fillna(df.mean(numeric_only=True), inplace=True) # Verify the results print(df.dtypes) print(df.head())
Quick Notes
errors='coerce'is the game-changer here—it prevents the function from throwing errors when it hits non-numeric strings.- Cleaning the quotes first is critical because
"notpresent"(with quotes) is treated as a different string thannotpresent(without), and the former won't be recognized as a missing value indicator otherwise. - Choosing to fill missing values instead of dropping rows will help you keep more data for analysis—just make sure the fill method makes sense for your specific dataset!
内容的提问来源于stack exchange,提问作者Hussain
相关产品推荐
相关产品推荐

