You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计DataFrame中diff字段的版本变更数量并分类?

问题场景

我的DataFrame包含diff字段,该字段是JSON格式的变更记录,代码里的结构如下:

"diff": {
    "paths": {
      "added": [
        "/login/token"
      ]
    },
    "endpoints": {
      "added": [
        {
          "method": "POST",
          "path": "/login/token"
        }
      ]
    }
  }

导出到CSV后,该字段呈现为字符串格式:

{'paths': {'added': ['/login/token']}, 'endpoints': {'added': [{'method': 'POST', 'path': '/login/token'}]}}

我需要实现以下需求:

  1. 按commitdate和info_version分组,统计每组内发生变更的字段/子字段总数
  2. 对变更进行新增、删除、修改分类统计

处理思路与步骤

1. 解析diff字段,统一结构化数据

首先要把CSV里的字典字符串转换成可操作的Python字典:

  • 用ast.literal_eval()解析CSV中的字符串(避免使用eval(),防止安全风险)
  • 如果是直接从JSON源加载的DataFrame,可直接用pd.json_normalize()展开字段

示例代码:

import ast
import pandas as pd

# 读取CSV后解析diff字段
df['diff'] = df['diff'].apply(ast.literal_eval)

2. 提取变更类型与计数

遍历diff字典的每个顶层键(比如paths、endpoints、info等),再遍历每个键下的操作类型(added/removed/modified),统计每个操作对应的条目数:

自定义处理单条diff记录的函数:

def count_changes(diff_dict):
    change_counts = {'added': 0, 'removed': 0, 'modified': 0, 'total': 0}
    for category in diff_dict.values():
        for op, items in category.items():
            if op in change_counts:
                count = len(items)
                change_counts[op] += count
                change_counts['total'] += count
    return pd.Series(change_counts)

# 生成统计列
df[['added', 'removed', 'modified', 'total_changes']] = df['diff'].apply(count_changes)

3. 按commitdate和info_version分组统计

用分组聚合函数,对每个分组的变更数求和:

grouped_stats = df.groupby(['commitdate', 'info_version']).agg(
    total_added=('added', 'sum'),
    total_removed=('removed', 'sum'),
    total_modified=('modified', 'sum'),
    total_changes=('total_changes', 'sum')
).reset_index()

4. 处理嵌套子字段的特殊情况

如果遇到info这类包含子字段变更的场景(比如title、description修改),需要根据实际结构调整统计逻辑:

  • 若modified下是字典(如{'title': {'old': '旧值', 'new': '新值'}}),则每个键对应一次修改,计数加1
  • 调整自定义函数的判断逻辑:
def count_changes(diff_dict):
    change_counts = {'added': 0, 'removed': 0, 'modified': 0, 'total': 0}
    for category in diff_dict.values():
        for op, items in category.items():
            if op == 'modified':
                # 处理子字段修改:每个字典的键对应一次变更
                if isinstance(items, dict):
                    count = len(items.keys())
                elif isinstance(items, list):
                    # 若modified是列表,每个元素算一次变更
                    count = len(items)
                change_counts[op] += count
                change_counts['total'] += count
            elif op in change_counts:
                count = len(items)
                change_counts[op] += count
                change_counts['total'] += count
    return pd.Series(change_counts)

内容的提问来源于stack exchange,提问作者Brie MerryWeather

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 19:45:31