You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame的section_name列提取内容并正确填充section_id列?

问题:DataFrame中section_id列的规则填充实现

原始数据与预期结果

原始DataFrame(含空白行)结构如下:

section_id  section_name
            1.Test Summary9
            1.1.Synopsis9
            1.2.Schema12
            1.3.1.Test Period  I - Screening13
            1.3.2.Period II - obes-Treatment 15
            Synopsis

            Test Period  I - Screening

需要按照以下规则填充section_id列:

  • 若section_name以section_id格式开头(如1.Test Summary9),直接提取该section_id填充;
  • 若section_name是已有带ID名称的简化版(如Synopsis对应1.1.Synopsis9),填充对应的section_id;
  • 空白行保持空值不处理。

最终预期结果:

section_id  section_name
1           1.Test Summary9
1.1         1.1.Synopsis9
1.2         1.2.Schema12
1.3.1       1.3.1.Test Period  I - Screening13
1.3.2       1.3.2.Period II - obes-Treatment 15
1.1         Synopsis
1.3.1       Test Period  I - Screening

用户尝试的代码

import pandas as pd

data = {
    'section_name': [
        '1.Test Summary9',
        '1.1.Synopsis9',
        '1.2.Schema12',
        '1.3.1.Test Period  I - Screening13',
        '1.3.2.Period II - obes-Treatment 15',
        'Synopsis',
        'Test Period  I - Screening'
    ]
}

df = pd.DataFrame(data)

def extract_section_id(section_name, current_section_id):
    if section_name.startswith(current_section_id):
        return current_section_id
    else:
        return section_name.split('.')[0]

current_section_id = ''
section_ids = []

for index, row in df.iterrows():
    section_name = row['section_name'].strip()
    if section_name != '':
        section_id = extract_section_id(section_name, current_section_id)
        current_section_id = section_id
    else:
        section_id = ''
    section_ids.append(section_id)

df['section_id'] = section_ids

print(df)

问题分析与最优实现

原代码的问题

原代码逻辑无法处理简化版名称匹配的场景,比如Synopsis无法对应到1.1——它既不以前面的1.3.2开头,也不能通过split('.')[0]提取到正确ID。

实现思路

  1. 先遍历数据,提取所有带section_id的条目,建立名称关键词与section_id的映射字典:
    • 对每个带ID的section_name,用正则提取完整层级的section_id(如1.1.Synopsis9提取1.1);
    • 提取名称核心关键词:去掉开头的ID和末尾的数字后缀,得到可匹配的纯名称(如1.1.Synopsis9处理为Synopsis);
  2. 再次遍历数据,按规则填充section_id:
    • 空白行直接留空;
    • 若名称以ID开头,直接提取ID;
    • 若名称是映射字典中的关键词,返回对应的ID。

完整实现代码

import pandas as pd
import re

# 构造包含空白行的原始数据
data = {
    'section_name': [
        '1.Test Summary9',
        '1.1.Synopsis9',
        '1.2.Schema12',
        '1.3.1.Test Period  I - Screening13',
        '1.3.2.Period II - obes-Treatment 15',
        'Synopsis',
        '',
        'Test Period  I - Screening'
    ]
}

df = pd.DataFrame(data)

# 第一步:构建名称关键词与section_id的映射
name_id_map = {}

for idx, row in df.iterrows():
    name = row['section_name'].strip()
    if not name:
        continue
    # 匹配层级格式的section_id(如1.、1.1.、1.3.1.)
    id_match = re.match(r'^(\d+(?:\.\d+)*)\.', name)
    if id_match:
        section_id = id_match.group(1)
        # 提取名称核心关键词:去掉ID前缀和末尾数字
        keyword = re.sub(r'^' + re.escape(section_id) + r'\.', '', name)
        keyword = re.sub(r'\d+$', '', keyword).strip()
        name_id_map[keyword] = section_id

# 第二步:填充section_id列
def fill_section_id(name):
    name_stripped = name.strip()
    if not name_stripped:
        return ''
    # 优先检查是否以ID开头
    id_match = re.match(r'^(\d+(?:\.\d+)*)\.', name_stripped)
    if id_match:
        return id_match.group(1)
    # 匹配映射字典中的关键词
    if name_stripped in name_id_map:
        return name_id_map[name_stripped]
    # 无匹配时返回空(可根据需求调整)
    return ''

df['section_id'] = df['section_name'].apply(fill_section_id)
# 调整列顺序与预期结果一致
df = df[['section_id', 'section_name']]

print(df)

运行结果

section_id                          section_name
0          1                      1.Test Summary9
1        1.1                        1.1.Synopsis9
2        1.2                          1.2.Schema12
3      1.3.1  1.3.1.Test Period  I - Screening13
4      1.3.2    1.3.2.Period II - obes-Treatment 15
5        1.1                              Synopsis
6                                               
7      1.3.1                Test Period  I - Screening

内容的提问来源于stack exchange,提问作者ista120

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 06:42:49