使用Pandas to_datetime转换日期列全返回NaT的问题求助
问题修复方案
核心问题:pd.to_datetime的格式参数错误
你的Month列格式是Oct-22、Sep-22这种英文月份缩写+两位年份,但你用了%m-%y格式——这个格式对应的是数字月份+两位年份(比如10-22),格式不匹配自然会返回全NaT。
修复步骤:
- 修正格式参数:把
format='%m-%y'改成format='%b-%y',%b是pandas识别英文月份缩写的格式符。 - 检查PEA列NaN问题:你读CSV时加了
dtype={'a': str},如果你的数据源里根本没有a列,这个参数可能会引发数据读取异常,导致其他列(比如PEA)丢失或变成NaN,建议直接删掉这个参数,或者替换成实际存在的列名。 - 可选:适配系统语言环境:如果你的系统默认是中文locale,pandas可能无法识别英文月份缩写,需要提前设置英文locale:
- Windows系统:
import locale locale.setlocale(locale.LC_TIME, 'English_US.1252') - Linux/macOS系统:
import locale locale.setlocale(locale.LC_TIME, 'en_US.UTF-8')
- Windows系统:
修改后的完整代码:
import numpy as np import pandas as pd # 可选:设置英文locale(如果系统是中文环境) # import locale # locale.setlocale(locale.LC_TIME, 'English_US.1252') # Windows用这个 # locale.setlocale(locale.LC_TIME, 'en_US.UTF-8') # Linux/macOS用这个 inputpath='HistoricalPrices' # 去掉不必要的dtype参数,避免读取异常 dataset=pd.read_csv(inputpath, sep=',', low_memory=False) print(dataset.head()) dataset = dataset.iloc[::-1].reset_index(drop=True) # 修正format参数为%b-%y dataset['Month']=pd.to_datetime(dataset['Month'], errors='coerce', format='%b-%y') print(dataset['Month'].dtypes) print(dataset)
验证说明:
运行修改后的代码后,Month列应该会正确转换为datetime64[ns]类型,且不再是全NaT;PEA列也会恢复正常数值(只要数据源本身该列有有效值)。
内容的提问来源于stack exchange,提问作者EcoHealthGuy
相关产品推荐
相关产品推荐

