You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于CSV DictReader字典列表筛选唯一ID对应最新日期条目

解决保留唯一ID最新日期行的问题

没问题,我来帮你把这个需求转换成可运行的Python代码!核心思路其实很清晰:我们需要先把字符串格式的日期转换成能直接比较的datetime对象,然后用ID作为键,遍历每一行数据,不断更新每个ID对应的最新日期条目。

直接上代码,结合你的示例来写:

from datetime import datetime

# 你的原始数据
rows = [
    {"id":"123","date":"1/1/18","foo":"bar"},
    {"id":"123","date":"2/2/18", "foo":"baz"}
]

# 用来存储最终结果的字典,键是id,值是最新的行数据
latest_rows = {}

# 遍历每一行数据
for row in rows:
    # 把字符串日期转换成datetime对象,注意这里的格式是月/日/年(%m/%d/%y)
    # 如果你的日期格式不同,需要调整这里的格式字符串
    current_date = datetime.strptime(row["date"], "%m/%d/%y")
    
    # 检查当前id是否已经在latest_rows里
    if row["id"] not in latest_rows:
        # 如果不在,直接添加进去,用copy避免后续修改影响原数据
        latest_rows[row["id"]] = row.copy()
        # 临时存储转换后的日期对象,方便后续比较
        latest_rows[row["id"]]["_date_obj"] = current_date
    else:
        # 如果已经存在,对比当前日期和已存储的日期
        existing_date = latest_rows[row["id"]]["_date_obj"]
        if current_date > existing_date:
            # 当前日期更新,替换成新的行数据
            latest_rows[row["id"]] = row.copy()
            latest_rows[row["id"]]["_date_obj"] = current_date

# 清理临时添加的_date_obj字段,得到干净的结果
result = {
    id: {key: val for key, val in item.items() if key != "_date_obj"}
    for id, item in latest_rows.items()
}

# 输出结果
print(result)
# 输出:{'123': {'id': '123', 'date': '2/2/18', 'foo': 'baz'}}

关键细节说明:

  • 日期转换的必要性:直接比较字符串日期会出错(比如"10/10/18"和"2/2/18"按字符串比较时,"2/2/18"会被判定为更大),转换成datetime对象才能正确判断先后顺序。
  • 格式字符串调整:如果你的CSV日期格式不是月/日/年,比如是年-月-日,只需要把strptime的格式参数改成"%Y-%m-%d"即可,根据实际格式灵活调整。
  • 字典的更新逻辑:每次遍历都检查当前ID是否已存在,不存在就新增,存在就对比日期后决定是否替换,确保每个ID始终保留最新的条目。

内容的提问来源于stack exchange,提问作者foobarbaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:35:48