You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何转换含未补零、空格的不规范字符串日期列?

Fixing Messy, Inconsistent Date Formats in Python/Pandas

Got it, let's work through this date formatting headache you're facing. The problem here is twofold: your dates have random extra spacing (like 2/9 /2018 with that stray space after the day) and single-digit months/days that don't play nice with strict strptime formatting rules. Here's how to fix it step by step:

1. First, Clean Up the Messy Strings

Before parsing, we need to eliminate those random spaces that are throwing off the format. Using pandas, we can strip all whitespace from each date string with a simple regex replace:

import pandas as pd

# Remove all spaces from the date column
df['start'] = df['start'].str.replace('\s+', '', regex=True)

This turns 2/9 /2018 into 2/9/2018, which is already way easier to work with.

2. Use Flexible Date Parsing (No Strict Formats Needed)

Your original strptime call failed because it was expecting a %m %d %Y format (month-day-year separated by spaces), but your dates use slashes. Even if you switched to %m/%d/%Y, strptime is strict—it requires two-digit months/days, so 1/26/2018 would still throw an error because the month is a single digit.

Instead, use one of these more forgiving approaches:

Option A: Pandas' Built-in to_datetime (Simplest)

Pandas has a fantastic to_datetime function that automatically handles most common date variations, including single-digit months/days and (funnily enough) even the original spaced-out dates without pre-cleaning:

# Parse the date column directly, even with leftover spaces
df['start_datetime'] = pd.to_datetime(df['start'], errors='coerce')

The errors='coerce' flag will turn any unparseable dates into NaT (Not a Time) instead of crashing your code—super useful for messy real-world data.

Option B: dateutil.parser (More Control)

If you need more customization, the dateutil library's parser can auto-detect almost any date format. First install it if you haven't (pip install python-dateutil), then use it with apply:

from dateutil import parser

# Parse cleaned date strings
df['start_datetime'] = df['start'].apply(lambda x: parser.parse(x))

This works great for edge cases that pandas might miss, though for your specific example, pandas alone is more than enough.

Why Your Original Code Failed

Just to clarify: datetime.datetime.strptime(df['start'][1], '%m %d %Y') was looking for a date like 01 26 2018 (month, day, year separated by spaces), but your data uses slashes (1/26/2018). Even if you corrected the format to %m/%d/%Y, strptime requires two-digit months/days—so 1/26/2018 would still fail because the month is a single digit. Flexible parsers avoid this strictness entirely.

内容的提问来源于stack exchange,提问作者mitch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:32:16