You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas替换Test case ID为连续序号

解决思路与实现代码

步骤1:读取并预处理原始数据

原始数据是带重复表头的固定宽度文本,先读取所有行,跳过重复表头,提取有效数据行:

import pandas as pd

# 读取原始文本文件(替换为你的文件路径)
with open('test_cases.txt', 'r') as f:
    lines = [line.strip() for line in f if line.strip()]

# 过滤掉重复的表头行,同时记录每个测试用例的起始位置
data_lines = []
case_starts = []
for idx, line in enumerate(lines):
    if line.startswith('Test case ID'):
        continue
    if line.startswith('id:'):
        case_starts.append(len(data_lines))
    data_lines.append(line)

# 用固定宽度格式解析数据
df = pd.read_fwf(
    pd.io.common.StringIO('\n'.join(data_lines)),
    colspecs='infer',
    names=['Test case ID', 'Title', 'Step Action', 'Expected Result']
)

步骤2:生成连续序号并填充

通过标记的测试用例起始位置,给每个测试用例分配连续ID,再填充同一测试用例的所有行:

# 创建测试用例分组标识
df['case_group'] = 0
for i in range(len(case_starts)):
    start_idx = case_starts[i]
    end_idx = case_starts[i+1] if i+1 < len(case_starts) else len(df)
    df.loc[start_idx:end_idx-1, 'case_group'] = i+1

# 替换原ID为连续序号,并向下填充空值
df['Test case ID'] = df['case_group']
df['Test case ID'] = df['Test case ID'].ffill()

# 移除临时分组列
df = df.drop('case_group', axis=1)

步骤3:还原格式输出

将处理后的数据按照原格式输出,包含重复的表头:

output_content = []
current_case = 1

for idx, row in df.iterrows():
    # 每个测试用例开头添加表头
    if row['Test case ID'] == current_case and (idx == 0 or df.loc[idx-1, 'Test case ID'] != current_case):
        output_content.append('Test case ID    Title    Step Action    Expected Result')
        current_case += 1
    # 格式化每行数据,保持对齐
    line = f"{str(row['Test case ID']).ljust(16)}" \
           f"{str(row['Title']).ljust(10) if pd.notna(row['Title']) else ''.ljust(10)}" \
           f"{str(row['Step Action']).ljust(14)}" \
           f"{str(row['Expected Result'])}"
    output_content.append(line)

# 写入文件或打印结果
with open('processed_test_cases.txt', 'w') as f:
    f.write('\n'.join(output_content) + '\n')

# 直接打印预览
print('\n'.join(output_content))

关键说明

  • 用read_fwf适配固定宽度的文本格式,自动识别列边界;
  • 通过标记测试用例起始行实现分组,ffill()快速填充同一测试用例的所有行ID;
  • 输出时还原原有的重复表头和对齐格式,保证结果与需求完全匹配。

内容的提问来源于stack exchange,提问作者Gia Huy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 07:35:39