You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Logistic Regression模型Streamlit部署时遇数值类型兼容错误求助

乳腺癌分类模型Streamlit部署报错排查:ValueError: dtype='numeric' is not compatible with arrays of bytes/strings

我用Logistic Regression训练了乳腺癌分类模型,用pickle保存后通过Streamlit部署,测试用户输入数据时出现以下错误:

ValueError: dtype='numeric' is not compatible with arrays of bytes/strings.
Convert your data to numeric values explicitly instead.

建模环节完整代码

import numpy as np
import pandas as pd
import sklearn.datasets
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score

breast_cancer_dataset = sklearn.datasets.load_breast_cancer()

d_frame = pd.DataFrame(breast_cancer_dataset.data, columns = breast_cancer_dataset.feature_names)
d_frame['label'] = breast_cancer_dataset.target
d_frame.isnull().sum()
d_frame.describe()
d_frame['label'].value_counts()
d_frame.groupby('label').mean()

X = d_frame.drop(columns = 'label', axis = 1)
Y = d_frame['label']

X_train, X_test, Y_train, Y_test = train_test_split(X, Y, test_size = 0.2, random_state = 2)

model = LogisticRegression()
model.fit(X_train, Y_train)

X_train_prediction = model.predict(X_train)
training_data_acc = accuracy_score(Y_train, X_train_prediction)

X_test_prediction = model.predict(X_test)
testing_data_acc = accuracy_score(Y_test, X_test_prediction)

ip_data = (13.54,14.36,87.46,566.3,0.09779,0.08129,0.06664,0.04781,0.1885,0.05766,0.2699,0.7886,2.058,23.56,0.008462,0.0146,0.02387,0.01315,0.0198,0.0023,15.11,19.26,99.7,711.2,0.144,0.1773,0.239,0.1288,0.2977,0.07259)
ip_as_numpy_arr = np.asarray(ip_data)
ip_data_reshaped = ip_as_numpy_arr.reshape(1,-1)

prediction = model.predict(ip_data_reshaped)
if(prediction[0] == 0):
    print("The Breast Cancer is Malignant")
else:
    print("The Breast Cancer is Benign")

import pickle

filename = 'trained_model.sav'
pickle.dump(model, open(filename, 'wb'))

loaded_model = pickle.load(open('trained_model.sav', 'rb'))

ip_data = (13.54,14.36,87.46,566.3,0.09779,0.08129,0.06664,0.04781,0.1885,0.05766,0.2699,0.7886,2.058,23.56,0.008462,0.0146,0.02387,0.01315,0.0198,0.0023,15.11,19.26,99.7,711.2,0.144,0.1773,0.239,0.1288,0.2977,0.07259)
ip_as_numpy_arr = np.asarray(ip_data)
ip_data_reshaped = ip_as_numpy_arr.reshape(1,-1)

prediction = loaded_model.predict(ip_data_reshaped)
if(prediction[0] == 0):
    print("The Breast Cancer is Malignant")
else:
    print("The Breast Cancer is Benign")

错误原因

核心问题是Streamlit获取的用户输入默认是字符串类型,而模型训练时使用的是数值型数据,直接将字符串数据传入模型预测会触发类型不兼容错误。本地测试用的是硬编码的数值元组,所以没有问题,但Streamlit输入框返回的字符串无法被模型识别。

解决办法

1. 显式转换用户输入为数值类型

在Streamlit代码中,获取输入后必须将每个值转换为float类型,再整理成模型可接受的数组格式:

import streamlit as st
import numpy as np
import pickle

# 加载训练好的模型
loaded_model = pickle.load(open('trained_model.sav', 'rb'))

# 生成30个特征输入框(对应乳腺癌数据集的30个特征)
feature_list = []
for idx in range(30):
    input_val = st.text_input(f"特征{idx+1}")
    feature_list.append(input_val)

if st.button("开始预测"):
    try:
        # 转换所有输入为数值类型,整理成模型需要的形状
        numeric_inputs = np.array([float(val) for val in feature_list]).reshape(1, -1)
        prediction = loaded_model.predict(numeric_inputs)
        
        # 输出预测结果
        if prediction[0] == 0:
            st.write("乳腺癌为恶性(Malignant)")
        else:
            st.write("乳腺癌为良性(Benign)")
    except ValueError as e:
        st.error(f"输入错误:{e},请确保所有输入都是有效数字")

2. 优化输入方式(可选)

用st.number_input替代st.text_input,直接获取数值类型输入,减少转换步骤:

feature_list = []
for idx in range(30):
    input_val = st.number_input(f"特征{idx+1}", value=0.0)
    feature_list.append(input_val)

# 直接转换为数组即可
numeric_inputs = np.array(feature_list).reshape(1, -1)

3. 额外验证

确保转换后的输入数组维度为(1, 30),和训练时的特征数量一致,避免维度不匹配的额外错误。

内容的提问来源于stack exchange,提问作者Raw'sOfficial

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 11:50:06