You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python DataFrame处理:移除列表方括号并筛选指定支付方式的患者ID

解决方案:处理DataFrame中的列表格式列并筛选指定支付方式的患者

首先,咱们先把你的示例数据用pandas构造出来,方便后续操作:

import pandas as pd

data = {
    'PatientId': ['PAT10000', 'PAT10001', 'PAT10002', 'PAT10003', 'PAT10004'],
    'Payor': [['Cash', 'Britam'], ['Madison', 'Cash'], ['Cash'], ['Cash', 'Madison', 'Resolution'], ['CIC Corporate', 'Cash']]
}
df = pd.DataFrame(data)

1. 移除Payor列的方括号,将列表转为字符串格式

你的Payor列存的是列表,要去掉方括号最优雅的方式是把列表元素用逗号(或其他分隔符)拼接成字符串,这样既去掉了括号,格式也更规整:

# 将列表转为逗号分隔的字符串
df['Payor'] = df['Payor'].apply(', '.join)

处理后的DataFrame会变成这样:

PatientId Payor
0 PAT10000 Cash, Britam
1 PAT10001 Madison, Cash
2 PAT10002 Cash
3 PAT10003 Cash, Madison, Resolution
4 PAT10004 CIC Corporate, Cash

如果只是单纯想去掉方括号(不推荐,因为如果列表元素有特殊字符可能出问题),也可以用字符串替换:

df['Payor'] = df['Payor'].astype(str).str.replace(r'[\[\]]', '', regex=True)

但还是推荐用join的方式,更安全可靠。


2. 筛选至少使用过指定支付方式(如Madison)的患者ID

这里分两种情况,如果你还没把Payor列转为字符串(还是列表格式),可以直接判断元素是否在列表中:

target_payor = 'Madison'
# 筛选Payor列表中包含目标支付方式的行,提取PatientId
filtered_patients = df[df['Payor'].apply(lambda x: target_payor in x)]['PatientId']

如果已经把Payor转为字符串了,就用字符串包含判断:

filtered_patients = df[df['Payor'].str.contains(target_payor)]['PatientId']

两种方式得到的结果都是:

1    PAT10001
3    PAT10003
Name: PatientId, dtype: object

合并操作:先筛选再处理格式(更高效)

如果你的数据量很大,先筛选出符合条件的行再处理Payor的格式会更高效,避免处理所有数据:

target_payor = 'Madison'
# 先筛选
filtered_df = df[df['Payor'].apply(lambda x: target_payor in x)]
# 再处理Payor格式
filtered_df['Payor'] = filtered_df['Payor'].apply(', '.join)
# 提取PatientId
patient_ids = filtered_df['PatientId'].tolist()  # 转成列表方便使用

这样得到的patient_ids就是['PAT10001', 'PAT10003'],完全符合需求。

内容的提问来源于stack exchange,提问作者BoredGeek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:49:27