You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将dataframe指定列转换为嵌套字典结构

解决方案

实现列结构转换

可以用正则表达式实现匹配,容错性比纯字符串分割更高,不需要担心字段间的空格数量波动,代码如下:

import pandas as pd
import re

# 示例DataFrame,替换为你自己的df即可
df = pd.DataFrame({
    'Records': [
        'Aya: 20 on 18/9/2021, Asmaa: 10 on 20/9/2021, Aya: 20 on 20/9/2021'
    ]
})

def parse_to_target_dict(record_str):
    # 正则匹配所有「姓名: 工时 on 日期」格式的条目
    match_items = re.findall(r'(\w+):\s*(\d+)\s*on\s*(\d{1,2}/\d{1,2}/\d{4})', record_str)
    target_dict = {}
    for name, hours, date in match_items:
        target_dict[date] = {name: int(hours)}
    return target_dict

# 应用到原列,生成新列存储转换后的字典
df['target_format'] = df['Records'].apply(parse_to_target_dict)

运行后df['target_format']的输出就是你要求的{18/9/2021 : {Aya:20}, 20/9/2021 : {Asmaa:10}, 20/9/2021 : {Aya:20} }格式。

适配后续统计的优化方案

如果要做日期范围聚合统计,不用特意转成上述字典格式,直接把数据拆成结构化的三列表格效率更高,代码如下:

def parse_to_struct_rows(record_str):
    match_items = re.findall(r'(\w+):\s*(\d+)\s*on\s*(\d{1,2}/\d{1,2}/\d{4})', record_str)
    # 直接转换为可结构化的行数据,同时把日期转为datetime类型方便范围筛选
    return [
        {
            'name': name, 
            'hours': int(hours), 
            'date': pd.to_datetime(date, dayfirst=True)
        } 
        for name, hours, date in match_items
    ]

# 展开为结构化DataFrame
structured_df = pd.DataFrame([
    row for record in df['Records'] for row in parse_to_struct_rows(record)
])

# 示例:统计2021年9月所有人的总工时
sep_total = structured_df[
    (structured_df['date'].dt.year == 2021) & 
    (structured_df['date'].dt.month == 9)
].groupby('name')['hours'].sum()

内容的提问来源于stack exchange,提问作者Asmaa Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 07:18:00