解决TypeError: Could not convert to numeric及Year列转日期问题
问题:LSTM代码触发
TypeError: Could not convert to numeric错误 运行LSTM相关Python代码时触发TypeError: Could not convert to numeric错误,提示无法将Year列中形如['1/1/20121/2/20121/3/201212/31/2023']的内容转换为数值类型。
代码片段
file_path = 'TEST3.csv' # Update the path to your CSV file columns = ['Year','T', 'TM', 'Tm', 'PP', 'Yields_Blé_dur'] n_steps_in, n_steps_out = 3, 1 test_set_years = 5 df = load_data(file_path) # Preprocess the data df_processed = preprocess_data_columns(df, columns) # Extract 'Year' for plotting purposes years = df_processed['Year'].values
错误栈信息
TypeError Traceback (most recent call last) Cell In[44], line 10 7 df = load_data(file_path) 9 # Preprocess the data ---> 10 df_processed = preprocess_data_columns(df, columns) 11 # Extract 'Year' for plotting purposes 12 years = df_processed['Year'].values Cell In[33], line 2 1 def preprocess_data_columns(df, columns): ----> 2 df = df[columns].fillna(df.mean()) #Fill NaN values with the mean of the column 3 return df File c:\Users\hemic\AppData\Local\Programs\Python\Python311\Lib\site-packages\pandas\core\frame.py:11335, in DataFrame.mean(self, axis, skipna, numeric_only, **kwargs) 11327 @doc(make_doc("mean", ndim=2)) 11328 def mean( 11329 self, (...) 11333 **kwargs, 11334 ): > 11335 result = super().mean(axis, skipna, numeric_only, **kwargs) 11336 if isinstance(result, Series): 11337 result = result.__finalize__(self, method="mean") ... -> 1678 raise TypeError(f"Could not convert {x} to numeric") 1679 try: 1680 x = x.astype(np.complex128) Type Error: Could not convert ['1/1/20121/2/20121/3/201212/31/2023'] to numeric.
问题分析
- 核心错误:
preprocess_data_columns函数中对非数值类型的Year列执行df.mean()操作,而Year列的内容是拼接错误的日期字符串(多个日期连在一起,如1/1/20121/2/2012),无法被转换为数值计算均值。 - 额外需求:需要将Year列转换为合法日期类型,同时提取年份信息用于后续绘图。
修复方案
1. 修复Year列的拼接错误并转换为日期类型
通过正则表达式拆分拼接的日期字符串,转换为标准日期格式,并提取单独的年份列:
import pandas as pd import re def fix_year_column(df): # 拆分拼接的日期字符串,匹配MM/DD/YYYY格式 df['Year'] = df['Year'].apply( lambda x: re.findall(r'\d{1,2}/\d{1,2}/\d{4}', str(x))[0] if pd.notna(x) else x ) # 转换为日期类型,无效值转为NaT df['Year'] = pd.to_datetime(df['Year'], format='%m/%d/%Y', errors='coerce') # 提取年份用于绘图 df['Year_Only'] = df['Year'].dt.year return df
2. 修改预处理函数,仅对数值列计算均值
调整preprocess_data_columns函数,避免对日期列执行数值操作:
def preprocess_data_columns(df, columns): df = df[columns].copy() # 筛选仅数值类型的列 numeric_cols = df.select_dtypes(include=['number']).columns # 仅对数值列用均值填充NaN df[numeric_cols] = df[numeric_cols].fillna(df[numeric_cols].mean()) return df
3. 整合完整代码流程
file_path = 'TEST3.csv' # Update the path to your CSV file columns = ['Year','T', 'TM', 'Tm', 'PP', 'Yields_Blé_dur'] n_steps_in, n_steps_out = 3, 1 test_set_years = 5 df = load_data(file_path) # 先修复Year列并转换为日期 df = fix_year_column(df) # 预处理数据(仅数值列填充均值) df_processed = preprocess_data_columns(df, columns) # 提取年份用于绘图 years = df_processed['Year_Only'].values
额外说明
- 如果Year列的拼接规则不是每日日期,需要根据实际数据调整正则表达式的匹配逻辑。
pd.to_datetime的errors='coerce'参数会将无法转换的字符串转为NaT,后续可根据需求处理这些无效值。
内容的提问来源于stack exchange,提问作者leone
相关产品推荐
相关产品推荐

