You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame的版本策略添加分类标签列?

问题描述

有如下结构的Pandas DataFrame:

import pandas as pd
df = pd.DataFrame(
    {'id': {0: 84, 1: 84, 2: 84, 3: 84, 4: 124, 5: 125},
     'info_version': {0: '1.1.0',1: 'alpha', 2: '7.20345.98', 3: '2', 4: '${git.build.version}', 5: 'version not set'}
    }
)

需要对info_version列的各类版本值按规则分类,生成包含Version标签列的新DataFrame,规则如下:

  • 属于指定列表的名称类值归为Names标签,列表:
    Names = ["develop", "draft", "genesis", "living", "main", "master", "next", "BETA", "DEV", "VERSION", "ALL"]
    
  • 属于指定列表的以7开头的语义化版本归为SemanticVer7标签,列表:
    SemanticVer7 = ['7.0', '7.0.0','7.1.0', '7.1.1', '7.10.3', '7.10.4', '7.10.5', '7.18.0']
    
  • 同时需要兼容其他自定义规则(如原代码中的版本号映射、未设置版本、时间戳等分类)

此前尝试np.select和np.where无法覆盖大量规则,简单apply函数仅能实现少量规则,需更高效可扩展的解决方案。

解决方案

采用规则字典+向量化判断的方式,既保证可扩展性,又比逐行apply效率更高。核心思路是将每类规则封装为判断条件与对应标签,按优先级依次匹配。

步骤1:定义规则集合

按优先级从高到低排序所有规则,每个规则包含布尔判断条件和对应标签:

import pandas as pd
import numpy as np

# 初始化DataFrame(补充测试用例)
df = pd.DataFrame(
    {'id': {0: 84, 1: 84, 2: 84, 3: 84, 4: 124, 5: 125, 6: 126, 7: 127},
     'info_version': {0: '1.1.0',1: 'alpha', 2: '7.20345.98', 3: '2', 4: '${git.build.version}', 5: 'version not set', 6: 'develop', 7: '7.1.0'}
    }
)

# 定义规则列表(优先级从上到下依次降低)
Names = ["develop", "draft", "genesis", "living", "main", "master", "next", "BETA", "DEV", "VERSION", "ALL"]
SemanticVer7 = ['7.0', '7.0.0','7.1.0', '7.1.1', '7.10.3', '7.10.4', '7.10.5', '7.18.0']

rules = [
    # 1. Names标签规则
    (df['info_version'].isin(Names), 'Names'),
    # 2. SemanticVer7标签规则
    (df['info_version'].isin(SemanticVer7), 'SemanticVer7'),
    # 3. 自定义精确匹配规则
    (df['info_version'] == '1.1.0', '1.1.0'),
    (df['info_version'] == 'v1', 'V1'),
    (df['info_version'] == '0', '0'),
    (df['info_version'] == '7.8.1', '7.8.1'),
    (df['info_version'] == '2', 'Version 2'),
    # 4. 特殊文本匹配
    (df['info_version'] == 'version not set', 'Unversioned'),
    # 5. 时间戳匹配(正则判断)
    (df['info_version'].str.match(r'^\d{4}-\d{2}-\d{2}$'), 'Timestamps'),
    (df['info_version'].str.match(r'^\d{4}-\d{2}-\d{2}T\d{2}:\d{2}:\d{2}Z$'), 'Timestamps with code'),
    # 6. 模板格式值匹配
    (df['info_version'].str.startswith('${') & df['info_version'].str.endswith('}'), 'TemplateVersion')
]

步骤2:用np.select批量匹配规则

np.select可接收多组条件与对应值,按顺序匹配第一个满足条件的规则,完美解决多规则覆盖问题:

# 拆分规则的条件和标签
conditions = [cond for cond, label in rules]
labels = [label for cond, label in rules]

# 生成Version列,未匹配任何规则的设为'Undefined'
df['Version'] = np.select(conditions, labels, default='Undefined')

最终结果

运行后得到的DataFrame:

idinfo_versionVersion
841.1.01.1.0
84alphaUndefined
847.20345.98Undefined
842Version 2
124${git.build.version}TemplateVersion
125version not setUnversioned
126developNames
1277.1.0SemanticVer7

扩展说明

  • 添加新规则只需在rules列表中按优先级插入新的(条件, 标签)元组,无需修改核心逻辑
  • 复杂判断可通过自定义函数生成布尔Series,例如:
    def is_custom_version(val):
        return val.startswith('v') and val[1:].isdigit()
    rules.append((df['info_version'].apply(is_custom_version), 'CustomV'))
    
  • 该方式基于向量化操作,比逐行apply效率更高,适合处理大规模数据

内容的提问来源于stack exchange,提问作者Brie MerryWeather

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 23:26:05