You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas如何拆分含/分隔符的test列 按值生成对应重复行

实现方案

核心思路说明

首先区分非test列和test列两类字段:

  • 非test列:拆分后的所有行取值和原行完全一致
  • test列:按/分割多值,不足拆分行数的位置填充X

根据你给出的文字描述和示例输出,存在两种对齐逻辑,可按需选择:

方案1:按拆分位置对齐(符合文字描述需求)

拆分行数取当前行所有test列拆分后的最大长度,每个test列按拆分后的顺序取对应位置的值,无对应位置则填X。

代码实现

import pandas as pd

# 示例数据
df = pd.DataFrame(data={
    'id': ["idx1"],
    'test_A': ["ABC", ],
    'test_B': ["X"],
    'test_C': ["ABC / XYZ"],
    'test_D': ["ABC / JKL / XYZ"]
})

# 区分列类型
non_test_cols = [col for col in df.columns if not col.startswith('test_')]
test_cols = [col for col in df.columns if col.startswith('test_')]

def process_row_by_position(row):
    split_data = {}
    max_row = 0
    # 拆分每个test列,计算最大拆分行数
    for col in test_cols:
        vals = [v.strip() for v in str(row[col]).split('/')]
        split_data[col] = vals
        if len(vals) > max_row:
            max_row = len(vals)
    # 补X到最大长度
    for col in test_cols:
        split_data[col] += ['X'] * (max_row - len(split_data[col]))
    # 构建结果行
    res_rows = []
    for i in range(max_row):
        new_row = {col: row[col] for col in non_test_cols}
        for col in test_cols:
            new_row[col] = split_data[col][i]
        res_rows.append(new_row)
    return pd.DataFrame(res_rows)

# 逐行处理合并结果
result = pd.concat(
    [process_row_by_position(row) for _, row in df.iterrows()],
    ignore_index=True
)
print(result)

该方案输出的test_D列值为["ABC", "JKL", "XYZ"]

方案2:按拆分值对齐(符合你给出的示例输出)

先提取当前行所有test列拆分后的所有唯一值,按出现顺序排序作为行维度,每个test列如果包含该值则取值,否则填X。

代码实现

import pandas as pd

# 示例数据
df = pd.DataFrame(data={
    'id': ["idx1"],
    'test_A': ["ABC", ],
    'test_B': ["X"],
    'test_C': ["ABC / XYZ"],
    'test_D': ["ABC / JKL / XYZ"]
})

# 区分列类型
non_test_cols = [col for col in df.columns if not col.startswith('test_')]
test_cols = [col for col in df.columns if col.startswith('test_')]

def process_row_by_value(row):
    split_data = {}
    all_values = []
    # 拆分每个test列,收集所有值
    for col in test_cols:
        vals = [v.strip() for v in str(row[col]).split('/')]
        split_data[col] = set(vals)
        all_values.extend(vals)
    # 去重保留出现顺序
    unique_values = list(dict.fromkeys(all_values))
    # 构建结果行
    res_rows = []
    for val in unique_values:
        new_row = {col: row[col] for col in non_test_cols}
        for col in test_cols:
            new_row[col] = val if val in split_data[col] else 'X'
        res_rows.append(new_row)
    return pd.DataFrame(res_rows)

# 逐行处理合并结果
result = pd.concat(
    [process_row_by_value(row) for _, row in df.iterrows()],
    ignore_index=True
)
print(result)

该方案输出完全匹配你给出的示例结果。

性能说明

数千行数据使用上述两种方案都可以流畅运行,无需额外优化。如果后续数据量提升到十万行以上,可以改用pandas向量化操作进一步提升处理速度。

内容的提问来源于stack exchange,提问作者Awans

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 19:48:02