You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas DataFrame如何从字典列表列提取组织名称

Pandas提取字符串格式字典列表中的组织名方案

问题原因

之前运行的代码返回全量NaN,核心原因是Organizations列存储的是字符串格式的字典列表,不是Python原生的列表/字典对象:直接用.str[0]只能取到字符串的第一个字符[,自然无法索引到organization字段。

需求样例

原始数据

IDOrganizations
1[{organization=Glaxosmithkline, character_offset=10512}, {organization=Vulpes Fund, character_offset=13845}]
2[{organization=Amazon, character_offset=14589}, {organization=Sinovac, character_offset=18923}]

期望输出

IDOrganizations
1Glaxosmithkline, Vulpes Fund
2Amazon, Sinovac

实现代码

方案1:正则直接提取(推荐,代码简洁效率高)

无需做格式转换,直接通过正则匹配所有organization=后的字段值,自动清理值前后多余空格,最后用英文逗号拼接:

import re
import pandas as pd

def extract_org(text):
    # 空值直接返回空字符串
    if pd.isna(text):
        return ''
    # 匹配organization=后到下一个逗号/右大括号之间的内容,去除前后空格
    orgs = [match.strip() for match in re.findall(r'organization=([^,}]+)', text)]
    return ', '.join(orgs)

# 生成新列
latin_combined['newOrg'] = latin_combined['organizations'].apply(extract_org)

方案2:格式转换后提取(适合需要同时取其他字段的场景)

如果后续还需要提取character_offset等其他字段,可以先把字符串转换为Python原生列表字典结构,再取值:

import ast
import pandas as pd

def parse_and_extract(text):
    if pd.isna(text):
        return ''
    # 将字典键值对的=替换为:,适配Python字典语法
    formatted_text = text.replace('=', ':')
    # 安全转换为原生列表
    org_items = ast.literal_eval(formatted_text)
    # 提取组织名并清理空格
    orgs = [item['organization'].strip() for item in org_items]
    return ', '.join(orgs)

latin_combined['newOrg'] = latin_combined['organizations'].apply(parse_and_extract)

验证结果

用提供的前5行样本测试,新列输出如下,无NaN值:

行索引newOrg值
0Vac, Health
1Store, Museum
2Mart, Rep
3Lodge, Hotel
4Airport, Landmark

注意事项

  • 两个方案都会自动清理组织名前后的多余空格(比如样本中 Vac、 Museum这类带前后空格的值会被自动修正),输出格式统一
  • 若组织名本身包含逗号,可调整正则匹配规则适配特殊格式

内容的提问来源于stack exchange,提问作者Anna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 00:12:53