如何在Pandas中截取'domain'后的目标内容生成新列?
提取字符串中'domain'后的目标编号
问题说明
存在格式固定的字段,示例值如下:
Non Compliance Risk Situation 1 - domain D-G101 Regulatory licenses and relations with regulators
需求是创建新列,仅保留'domain'后的D-G101内容。尝试以下代码后未达到预期,新列值为[]:
osi['col'] = osi.Name.str.split('domain').str.slice(start=7) # 预期得到'D-G101',实际返回[]
可行解决方案
方法1:拆分字符串后精准索引提取
利用split拆分后定位目标内容:
osi['col'] = osi.Name.str.split('domain').str[1].str.split().str[0]
str.split('domain').str[1]:获取'domain'之后的全部内容(示例中为' D-G101 Regulatory licenses...')str.split().str[0]:将上述内容按空格拆分后取第一个元素,即目标的D-G101
方法2:正则表达式匹配(更适配固定格式)
通过正则直接捕获'domain '后的非空格字符,稳定性更强:
# 方法2.1 使用apply结合re.search import re osi['col'] = osi.Name.apply(lambda x: re.search(r'domain (\S+)', x).group(1)) # 方法2.2 使用pandas内置的str.extract osi['col'] = osi.Name.str.extract(r'domain (\S+)', expand=False)
正则表达式domain (\S+)的作用是匹配'domain '后面的连续非空格字符,捕获组(\S+)会直接提取出我们需要的D-G101。
内容的提问来源于stack exchange,提问作者ramoss
相关产品推荐
相关产品推荐

