You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决TypeError: Could not convert to numeric及Year列转日期问题

问题:LSTM代码触发TypeError: Could not convert to numeric错误

运行LSTM相关Python代码时触发TypeError: Could not convert to numeric错误,提示无法将Year列中形如['1/1/20121/2/20121/3/201212/31/2023']的内容转换为数值类型。

代码片段

file_path = 'TEST3.csv'  # Update the path to your CSV file
columns = ['Year','T', 'TM', 'Tm', 'PP', 'Yields_Blé_dur']
n_steps_in, n_steps_out = 3, 1
test_set_years = 5


df = load_data(file_path)

# Preprocess the data
df_processed = preprocess_data_columns(df, columns)
# Extract 'Year' for plotting purposes
years = df_processed['Year'].values

错误栈信息

TypeError                                 Traceback (most recent call last)
Cell In[44], line 10
      7 df = load_data(file_path)
      9 # Preprocess the data
---> 10 df_processed = preprocess_data_columns(df, columns)
     11 # Extract 'Year' for plotting purposes
     12 years = df_processed['Year'].values

Cell In[33], line 2
      1 def preprocess_data_columns(df, columns):
----> 2     df = df[columns].fillna(df.mean())  #Fill NaN values with the mean of the column
      3     return df

File c:\Users\hemic\AppData\Local\Programs\Python\Python311\Lib\site-packages\pandas\core\frame.py:11335, in DataFrame.mean(self, axis, skipna, numeric_only, **kwargs)
  11327 @doc(make_doc("mean", ndim=2))
  11328 def mean(
  11329     self,
   (...)
  11333     **kwargs,
  11334 ):
> 11335     result = super().mean(axis, skipna, numeric_only, **kwargs)
  11336     if isinstance(result, Series):
  11337         result = result.__finalize__(self, method="mean")
...
-> 1678     raise TypeError(f"Could not convert {x} to numeric")
   1679 try:
   1680     x = x.astype(np.complex128)

Type Error: Could not convert ['1/1/20121/2/20121/3/201212/31/2023'] to numeric.

问题分析

  • 核心错误:preprocess_data_columns函数中对非数值类型的Year列执行df.mean()操作,而Year列的内容是拼接错误的日期字符串(多个日期连在一起,如1/1/20121/2/2012),无法被转换为数值计算均值。
  • 额外需求:需要将Year列转换为合法日期类型,同时提取年份信息用于后续绘图。

修复方案

1. 修复Year列的拼接错误并转换为日期类型

通过正则表达式拆分拼接的日期字符串,转换为标准日期格式,并提取单独的年份列:

import pandas as pd
import re

def fix_year_column(df):
    # 拆分拼接的日期字符串,匹配MM/DD/YYYY格式
    df['Year'] = df['Year'].apply(
        lambda x: re.findall(r'\d{1,2}/\d{1,2}/\d{4}', str(x))[0] if pd.notna(x) else x
    )
    # 转换为日期类型,无效值转为NaT
    df['Year'] = pd.to_datetime(df['Year'], format='%m/%d/%Y', errors='coerce')
    # 提取年份用于绘图
    df['Year_Only'] = df['Year'].dt.year
    return df

2. 修改预处理函数,仅对数值列计算均值

调整preprocess_data_columns函数,避免对日期列执行数值操作:

def preprocess_data_columns(df, columns):
    df = df[columns].copy()
    # 筛选仅数值类型的列
    numeric_cols = df.select_dtypes(include=['number']).columns
    # 仅对数值列用均值填充NaN
    df[numeric_cols] = df[numeric_cols].fillna(df[numeric_cols].mean())
    return df

3. 整合完整代码流程

file_path = 'TEST3.csv'  # Update the path to your CSV file
columns = ['Year','T', 'TM', 'Tm', 'PP', 'Yields_Blé_dur']
n_steps_in, n_steps_out = 3, 1
test_set_years = 5


df = load_data(file_path)

# 先修复Year列并转换为日期
df = fix_year_column(df)

# 预处理数据(仅数值列填充均值)
df_processed = preprocess_data_columns(df, columns)
# 提取年份用于绘图
years = df_processed['Year_Only'].values

额外说明

  • 如果Year列的拼接规则不是每日日期,需要根据实际数据调整正则表达式的匹配逻辑。
  • pd.to_datetime的errors='coerce'参数会将无法转换的字符串转为NaT,后续可根据需求处理这些无效值。

内容的提问来源于stack exchange,提问作者leone

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 10:31:37