You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取指定格式日期并按升序为对应行分配排序编号

实现方案(基于Python + Pandas)

核心实现逻辑

  • 用正则匹配行首符合月份缩写 日, 年 —格式的日期串
  • 将匹配到的日期串转换为datetime类型,无匹配的行赋值为NaT空时间类型
  • 筛选出所有时间不为空的行,按时间升序排序后分配从1开始的连续编号
  • 将编号映射回原表,未匹配到日期的行统一赋值为-1

完整可运行代码

import pandas as pd
from datetime import datetime

# 示例数据,可替换为你自己的数据集读取代码
data = {
    "Text": [
        "Jun 28, 2021 — Brendan Moore is p...",
        "Professor of Psychology at University",
        "Aug 24, 2019 — Chemistry (Nobel prize...",
        "by A Craig · 2019 · Cited by 1 — Authors. ...",
        "... 2020 | Volume 8 | Article 330Edited by:"
    ]
}
df = pd.DataFrame(data)

# 匹配行首日期并转换为datetime格式
date_pattern = r'^([A-Z][a-z]{2}) (\d{1,2}), (\d{4}) —'
def parse_date(text):
    match = pd.Series(text).str.extract(date_pattern, expand=False).iloc[0]
    if pd.notna(match[0]):
        date_str = f"{match[0]} {match[1]}, {match[2]}"
        return datetime.strptime(date_str, "%b %d, %Y")
    return pd.NaT

df['temp_date'] = df['Text'].apply(parse_date)

# 为有效日期行分配升序编号
valid_date_df = df[df['temp_date'].notna()].sort_values('temp_date', ascending=True).copy()
valid_date_df['Numbering (sort by date asc)'] = range(1, len(valid_date_df)+1)

# 映射编号回原表,无匹配项统一填-1
df = df.merge(valid_date_df[['Text', 'Numbering (sort by date asc)']], on='Text', how='left')
df['Numbering (sort by date asc)'] = df['Numbering (sort by date asc)'].fillna(-1).astype(int)

# 删除临时列输出结果
df = df.drop('temp_date', axis=1)
print(df)

输出效果

TextNumbering (sort by date asc)
Jun 28, 2021 — Brendan Moore is p...2
Professor of Psychology at University-1
Aug 24, 2019 — Chemistry (Nobel prize...1
by A Craig · 2019 · Cited by 1 — Authors. ...-1
... 2020Volume 8

内容的提问来源于stack exchange,提问作者LdM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 03:06:03