You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中提取带分隔符列的首个分段至新列

处理Pandas DataFrame中code列提取首个分段的方法

示例数据

import pandas as pd

raw_data = {'name': ['Willard Morris', 'Al Jennings', 'Omar Mullins', 'Spencer McDaniel'],
'code': ['01-02-11-55-00115','11-02-11-55-00445','test', '31-0t-11-55-00115'],
'favorite_color': ['blue', 'blue', 'yellow', 'green'],  
'grade': [88, 92, 95, 70]}
df = pd.DataFrame(raw_data)

解决方案

这里提供两种简洁的实现方式:

方法1:使用str.split()结合where()

通过分割字符串并条件赋值实现需求:

# 生成新列first_code
df['first_code'] = df['code'].str.split('-').str[0].where(df['code'].str.contains('-'), pd.NA)
  • str.split('-')将code列的每个值按-分割为字符串列表
  • str[0]提取列表的第一个元素
  • where()方法判断原code值是否包含-:满足则保留提取结果,不满足则设为pd.NA(Pandas标准空值)

方法2:使用正则表达式str.extract()

利用正则精准匹配并提取目标片段:

df['first_code'] = df['code'].str.extract(r'^([^-]+)(?=-)')
  • 正则表达式^([^-]+)(?=-)的含义:匹配字符串开头的一段非-字符,且这段字符之后紧跟-
  • 若code值不含-,则提取结果为NaN(空值)

最终结果

处理后的DataFrame中first_code列的值依次为:01、11、<NA>、31,完全符合预期。

内容的提问来源于stack exchange,提问作者Natali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 05:20:23