You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas中转换自定义格式日期字符串报错的解决方法

解决Pandas转换"日+英文月份缩写+两位年份"日期格式的问题

错误原因分析

你遇到的unconverted data remains错误和代码问题主要来自两点:

  1. 逻辑错误:用datetime.strptime处理单个行的日期值(df['StartDate'].iloc[i]),却直接赋值给整列df['StartDate'],既覆盖了正常数据,也没实现批量转换的需求——datetime.strptime本身是处理单个字符串的工具,不适合批量处理Pandas列。
  2. 解析匹配问题:
    • 部分日期字符串可能包含空格、换行符等多余字符,导致解析时出现未转换的剩余内容。
    • 若系统默认locale为中文,%b格式符无法识别Aug/Dec/Jan这类英文月份缩写,解析到月份部分就会卡住,触发错误。

解决方案

步骤1:先清理日期字符串

先去除日期字符串前后的多余字符:

# 单列清理
df['StartDate'] = df['StartDate'].str.strip()
# 多列批量清理
date_cols = ['StartDate', 'EndDate', 'CreateDate']
df[date_cols] = df[date_cols].apply(lambda x: x.str.strip())

步骤2:用Pandas批量转换(推荐)

使用pd.to_datetime批量处理整列,同时指定英文locale确保识别月份缩写:

import pandas as pd
import locale

# 设置英文locale(不同系统写法有差异:Linux/macOS用'en_US.UTF-8',Windows用'en-US')
locale.setlocale(locale.LC_TIME, 'en_US.UTF-8')

# 转换单列
df['StartDate'] = pd.to_datetime(df['StartDate'], format='%d%b%y', errors='coerce')

# 一次性转换3列
df[date_cols] = df[date_cols].apply(lambda col: pd.to_datetime(col, format='%d%b%y', errors='coerce'))
  • errors='coerce'会把解析失败的日期转为NaT(Pandas的空日期值),方便你用df[df['StartDate'].isna()]排查格式异常的行。

备选方案:自定义解析函数

如果设置locale遇到问题,可自定义函数精准解析:

from datetime import datetime
import pandas as pd
import locale

locale.setlocale(locale.LC_TIME, 'en_US.UTF-8')

def parse_date(date_str):
    try:
        dt = datetime.strptime(date_str, '%d%b%y')
        # 可选:手动调整两位年份的 century,比如把23转为2023而非1923
        # dt = dt.replace(year=dt.year + 2000 if dt.year < 50 else dt.year)
        return dt
    except ValueError:
        return pd.NaT

df['StartDate'] = df['StartDate'].apply(parse_date)

验证转换结果

转换完成后,可通过以下命令确认效果:

# 查看列的数据类型
print(df['StartDate'].dtype)
# 查看前几行转换结果
print(df['StartDate'].head())
# 定位转换失败的行
print(df[df['StartDate'].isna()])

内容的提问来源于stack exchange,提问作者iBeMeltin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 06:12:50