You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pd.to_numeric转换报错ValueError:无法解析‘notpresent’求解决方案

Fixing Numeric Conversion Error in Your Python Data Processing Code

Hey there, let's break down what's going wrong and fix that ValueError you're seeing.

What's Causing the Error?

The problem is straightforward: your bgr column contains non-numeric strings like "notpresent" (complete with extra double quotes!), and pd.to_numeric() can't parse these into floating-point numbers by default. That's exactly why you're getting the Unable to parse string error.

Step-by-Step Fix

Let's walk through the fixes one by one:

  1. Strip Unwanted Double Quotes
    First, we need to remove those extra double quotes from all string columns—they're blocking proper value conversion:
# Remove double quotes from any object-type columns
df = df.apply(lambda col: col.str.replace('"', '') if col.dtype == 'object' else col)
  1. Convert Columns Gracefully
    Use the errors='coerce' argument with pd.to_numeric()—this will turn any unparseable strings into NaN instead of crashing your code:
# Convert columns to float, turning bad values into NaN
df['bgr'] = pd.to_numeric(df['bgr'], errors='coerce', downcast='float')
df['wc'] = pd.to_numeric(df['wc'], errors='coerce', downcast='float')
df['rc'] = pd.to_numeric(df['rc'], errors='coerce', downcast='float')
  1. Handle Missing Values (Optional but Important)
    Now you'll have NaN values in your dataframe. You can either drop rows with missing values or fill them with a statistic like the mean/median (filling is usually better for retaining data):
# Option 1: Drop rows with any NaN values
df.dropna(axis=0, how='any', inplace=True)

# Option 2: Fill NaNs with column means
# df.fillna(df.mean(numeric_only=True), inplace=True)

Full Fixed Code

Here's your complete code with all the fixes applied:

import math
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib
from sklearn import preprocessing

# Load the dataset
df = pd.read_html('https://github.com/authman/DAT210x/blob/master/Module4/Datasets/kidney_disease.csv')[0]

# Drop the id column
df.drop('id', axis=1, inplace=True)

# Keep only the specified columns
df = df[['bgr', 'wc', 'rc']]

# Clean up double quotes from string columns
df = df.apply(lambda col: col.str.replace('"', '') if col.dtype == 'object' else col)

# Convert columns to numeric, coercing bad values to NaN
df['bgr'] = pd.to_numeric(df['bgr'], errors='coerce', downcast='float')
df['wc'] = pd.to_numeric(df['wc'], errors='coerce', downcast='float')
df['rc'] = pd.to_numeric(df['rc'], errors='coerce', downcast='float')

# Handle missing values (choose one option)
df.dropna(axis=0, how='any', inplace=True)
# df.fillna(df.mean(numeric_only=True), inplace=True)

# Verify the results
print(df.dtypes)
print(df.head())

Quick Notes

  • errors='coerce' is the game-changer here—it prevents the function from throwing errors when it hits non-numeric strings.
  • Cleaning the quotes first is critical because "notpresent" (with quotes) is treated as a different string than notpresent (without), and the former won't be recognized as a missing value indicator otherwise.
  • Choosing to fill missing values instead of dropping rows will help you keep more data for analysis—just make sure the fill method makes sense for your specific dataset!

内容的提问来源于stack exchange,提问作者Hussain

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:37:23