You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决ValueError: setting an array element with a sequence?蛋白序列入神经网络报错

Fixing "ValueError: setting an array element with a sequence" for Protein Sequence Neural Network Inputs

Alright, let's break down exactly what's going wrong here and walk through fixing your data pipeline to get your model working. That error pops up because your input data doesn't have a consistent shape—neural networks require fixed-dimensional tensors, but your protein sequences are likely varying lengths, and your current conversion code isn't addressing that properly.

1. First: Fix the Core Issue - Uniform Sequence Lengths

Protein sequences almost never have the same length, so the first step is to standardize them. You can either pad shorter sequences with a placeholder (like the rarely used amino acid 'X') or truncate longer ones to match the longest sequence in your dataset.

import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense, Embedding, Flatten
from tensorflow.keras.utils import to_categorical

# Load your dataset
df = pd.read_csv('/home/alpha/mk fyp/whole/DATASET2.csv', names=('X1','Y'), delimiter=',')

# Find the longest sequence length in your dataset
max_seq_length = df['X1'].str.len().max()

# Function to pad/truncate sequences to uniform length
def standardize_sequence(seq, max_len):
    if len(seq) >= max_len:
        return seq[:max_len]  # Truncate long sequences
    else:
        return seq.ljust(max_len, 'X')  # Pad short sequences with 'X'

# Apply standardization to all sequences
df['X1_standardized'] = df['X1'].apply(lambda x: standardize_sequence(x, max_seq_length))

2. Encode Protein Sequences Properly

Your original convert function isn't designed for protein sequences—amino acids are characters, not numeric values, so trying to cast them to floats just leaves you with raw character lists. Instead, use one of these two robust encoding methods:

This maps each unique amino acid to an integer, then uses an Embedding layer to turn those integers into meaningful dense vectors (great for capturing biological patterns in sequences):

# Collect all unique amino acids (including our placeholder 'X')
all_amino_acids = set(''.join(df['X1_standardized'].values))
# Create a mapping from amino acid to integer (start at 1 to reserve 0 for padding if needed)
aa_to_idx = {aa: idx+1 for idx, aa in enumerate(all_amino_acids)}
aa_to_idx['X'] = 0  # Assign 0 to our padding character

# Convert each standardized sequence to an array of integers
def seq_to_int_array(seq):
    return [aa_to_idx[aa] for aa in seq]

# Convert all sequences to a numpy array of integers
X = np.array(df['X1_standardized'].apply(seq_to_int_array).tolist())

Option 2: One-Hot Encoding (Good for Short Fixed-Length Sequences)

If your sequences are short, you can one-hot encode each amino acid, turning each sequence into a 2D tensor of shape (max_seq_length, number_of_unique_amino_acids):

from sklearn.preprocessing import OneHotEncoder

# Split each sequence into individual amino acid characters
split_seqs = df['X1_standardized'].apply(list).tolist()
# Flatten all amino acids to fit the encoder
all_flat_aas = [aa for seq in split_seqs for aa in seq]

# Initialize and fit the one-hot encoder
encoder = OneHotEncoder(sparse_output=False)
encoder.fit(np.array(all_flat_aas).reshape(-1, 1))

# One-hot encode each sequence
X = np.array([encoder.transform(np.array(seq).reshape(-1, 1)) for seq in split_seqs])
# Reshape to 2D if using a Dense-only model (flatten the sequence and one-hot dimensions)
X = X.reshape(X.shape[0], X.shape[1] * X.shape[2])

3. Fix Your Label Data

Your labels are a, ab, b—this is a multi-class classification problem, not binary. You need to encode these labels into one-hot vectors to match the model's output:

# Encode string labels to integers
label_encoder = LabelEncoder()
y_integer = label_encoder.fit_transform(df['Y'].values)
# Convert integers to one-hot vectors
Y = to_categorical(y_integer)

4. Adjust Your Neural Network Architecture

Your original model is designed for single numeric inputs, not sequences. Here's how to adjust it for your encoded data:

For Integer Encoding + Embedding Layer

def baseline_model():
    model = Sequential()
    # Embedding layer: maps integer indices to dense vectors
    model.add(Embedding(input_dim=len(aa_to_idx), output_dim=32, input_length=max_seq_length))
    model.add(Flatten())  # Flatten the 2D embedding output to 1D for Dense layers
    model.add(Dense(10, activation='relu', kernel_initializer='uniform'))
    # Output layer: use softmax for multi-class classification
    model.add(Dense(len(label_encoder.classes_), activation='softmax'))
    # Use categorical crossentropy for multi-class labels
    model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
    return model

For One-Hot Encoded Data

def baseline_model():
    model = Sequential()
    # Input shape is the flattened length of one-hot encoded sequences
    input_shape = (max_seq_length * len(encoder.categories_[0]),)
    model.add(Dense(10, activation='relu', kernel_initializer='uniform', input_shape=input_shape))
    model.add(Dense(len(label_encoder.classes_), activation='softmax'))
    model.compile(loss='categorical_crossentropy', optimizer='adam', metrics=['accuracy'])
    return model

5. Correct Training Code

Your original training code had a few critical mistakes—here's the fixed version:

# Split data into train/test sets correctly
X_train, X_test, y_train, y_test = train_test_split(X, Y, test_size=0.20, random_state=42)

# Build and train the model
model = baseline_model()
model.fit(X_train, y_train, validation_data=(X_test, y_test), epochs=10, batch_size=32)

# Evaluate the model
scores = model.evaluate(X_test, y_test, verbose=0)
print(f"Model Accuracy: {scores[1]*100:.2f}%")
print(f"Model Error: {(100 - scores[1]*100):.2f}%")

Why Your Original Code Failed

  • Your convert function didn't handle protein sequences correctly—it left you with variable-length character lists, which can't be converted into a uniform numpy array for the model.
  • The model's input_dim=1 was completely mismatched for sequence data.
  • You used binary_crossentropy for a multi-class problem, which is incompatible.
  • You passed raw X1 as validation data instead of your processed input tensor, causing shape mismatches.

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:11:01