如何基于df的Service列高效创建服务存在性布尔值新列?
如何高效为每个服务生成布尔值标记列(Pandas)
完全没问题,这在Pandas里有非常高效的实现方式,不用手动逐行判断,利用内置的字符串处理和向量化操作就能轻松搞定,既简洁又快速。
完整实现步骤
首先,我们先构造你的示例数据,然后一步步完成需求:
import pandas as pd # 1. 构造示例数据集 data = { 'Service': [ 'DoorDash, Grubhub / Seamless, UberEats, Postmates', 'DoorDash, UberEats, Caviar, Tock', 'DoorDash', 'None', 'Caviar, Tock', 'None', 'Tock', 'DoorDash, Grubhub / Seamless, UberEats, Postmates', 'Grubhub / Seamless, UberEats' ] } df = pd.DataFrame(data) # 2. 定义需要生成列的服务列表(包含你指定的所有服务,加上Tock以匹配预期输出) service_list = [ 'DoorDash', 'Grubhub / Seamless', 'UberEats', 'Caviar', 'Postmates', 'JustEat', 'Deliveroo', 'Foodora', 'Grab', 'Talabat', 'Tock' ] # 3. 高效生成布尔值列(推荐向量化方法,大数据集下性能更优) # 先处理'None'值为空白,避免干扰判断 df['Service'] = df['Service'].replace('None', '') # 分割服务字符串并生成哑变量,再按行求和得到存在标记 split_services = df['Service'].str.split(',\s*', expand=True) service_dummies = pd.get_dummies(split_services.stack()).sum(level=0) # 合并原数据和哑变量,补充服务列表中未出现的服务列(默认设为False) result_df = df.join(service_dummies.reindex(columns=service_list, fill_value=False)) # 如果需要把布尔值显式转为True/False(默认就是布尔类型,这步可选) result_df = result_df.astype({col: bool for col in service_list})
更直观的循环实现(适合新手理解)
如果你觉得上面的向量化方法有点绕,也可以用更直白的循环方式,逻辑清晰易懂:
# 处理'None'值 df['Service'] = df['Service'].replace('None', '') # 遍历每个服务,生成对应的布尔列 for service in service_list: # 检查该行是否包含当前服务,忽略大小写、空值设为False df[service] = df['Service'].str.contains(service, case=False, na=False)
输出结果
运行后得到的result_df就完全符合你的预期格式:
| Service | DoorDash | Grubhub / Seamless | UberEats | Caviar | Postmates | JustEat | Deliveroo | Foodora | Grab | Talabat | Tock |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DoorDash, Grubhub / Seamless, UberEats, Postmates | True | True | True | False | True | False | False | False | False | False | False |
| DoorDash, UberEats, Caviar, Tock | True | False | True | True | False | False | False | False | False | False | True |
| DoorDash | True | False | False | False | False | False | False | False | False | False | False |
| False | False | False | False | False | False | False | False | False | False | False | |
| Caviar, Tock | False | False | False | True | False | False | False | False | False | False | True |
| False | False | False | False | False | False | False | False | False | False | False | |
| Tock | False | False | False | False | False | False | False | False | False | False | True |
| DoorDash, Grubhub / Seamless, UberEats, Postmates | True | True | True | False | True | False | False | False | False | False | False |
| Grubhub / Seamless, UberEats | False | True | True | False | False | False | False | False | False | False | False |
关键注意点
- 处理
None:我们把字符串'None'替换为空,这样判断时会统一返回False,符合你的需求 - 大小写问题:
str.contains里的case=False可以避免大小写不一致导致的误判,如果你的数据是严格大小写匹配的,可以去掉这个参数 - 性能:向量化方法(第一种)比循环方法(第二种)在大数据集下快很多,推荐优先使用
内容的提问来源于stack exchange,提问作者Chris90
相关产品推荐
相关产品推荐

