You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame字符串列提取数据并生成对应新列?

解决方法

我们可以通过自定义解析函数处理Desc列的结构化文本,提取指定字段并生成新列,具体步骤如下:

1. 定义解析函数

编写函数将单个Desc字符串解析为键值对字典,清理字段名中的多余符号(如**、*),并提取冒号后的内容:

import pandas as pd

def parse_desc(desc_str):
    # 按双换行分割各字段条目
    fields = desc_str.split('\n\n')
    parsed_data = {}
    for field in fields:
        # 按冒号分割键和值(仅分割一次)
        if ': ' in field:
            key_raw, value = field.split(': ', 1)
            # 清理键名:移除特殊符号并去除首尾空格
            clean_key = key_raw.strip().replace('**', '').replace('*', '').strip()
            parsed_data[clean_key] = value.strip()
    return parsed_data

2. 应用解析函数并合并数据

将解析函数应用到Desc列,把生成的字典序列转为DataFrame,再与原DataFrame合并:

# 示例数据
data = {'sl no': [661, 662],
        'key': ['3484', '3483'],
        'id': [13592349, 13592490],
        'Sum': ['[E-1]', '[E-1]'],
        'Desc': [
              "**Title **: New_ind\n\n**Body **: Detection_error\n\n*respo_URL **: www.github.com\n\n**respo_status **: {yellow}","**Title **: New_ind2\n\n**Body **: import_error\n\n*respo_URL **: \n\n**respo_status **: {green}"]}

df = pd.DataFrame(data)

# 解析Desc列并生成新DataFrame
parsed_fields = df['Desc'].apply(parse_desc).apply(pd.Series)

# 合并原数据与解析后的字段(移除原Desc列)
final_df = pd.concat([df.drop('Desc', axis=1), parsed_fields], axis=1)

print(final_df)

输出结果

执行后会得到包含Title、Body、respo_URL、respo_status新列的DataFrame,示例输出如下:

sl no   key        id    Sum     Title           Body       respo_URL respo_status
0    661  3484  13592349  [E-1]    New_ind  Detection_error  www.github.com      {yellow}
1    662  3483  13592490  [E-1]  New_ind2    import_error                      {green}

内容的提问来源于stack exchange,提问作者Ashish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:32:19