You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python库拆分DataFrame文本列生成包含/排除列?

拆分DataFrame文本列获取包含与排除内容

思路

观察文本的固定模式:所有条目均以includes系列短语引出包含内容,随后通过but/however等转折词衔接doesn’t include引出排除内容。这种结构化文本无需复杂NLP模型,用正则表达式即可精准提取目标内容。

实现代码

1. 准备示例数据

import pandas as pd

data = {
    "column_description": [
        "this section includes: animals: cats and dogs and vegetables but doesn’t include: plants and fruits: coco",
        "this section includes the following: axis: x and y but doesn’t include: z, k and c",
        "this section includes notably: letters: a, b and c however it doesn’t include: y and letter: z"
    ]
}
df = pd.DataFrame(data)

2. 提取包含/排除内容

# 提取column_include:匹配includes后到转折词前的内容
df["column_include"] = df["column_description"].str.extract(
    r"includes.*?: (.*?) (but|however)", expand=False
)[0].str.strip()

# 提取column_exclude:匹配doesn’t include后的所有内容
df["column_exclude"] = df["column_description"].str.extract(
    r"doesn’t include: (.*)$", expand=False
).str.strip()

3. 查看结果

print(df[["column_include", "column_exclude"]])

输出结果:

column_include          column_exclude
0  animals: cats and dogs and vegetables  plants and fruits: coco
1                      axis: x and y           z, k and c
2           letters: a, b and c        y and letter: z

正则说明

  • includes.*?: (.*?) (but|however):
    • includes.*?: 非贪婪匹配从includes到最后一个冒号的内容(适配不同表述如includes the following:)
    • (.*?) 捕获冒号后到转折词前的目标文本
  • doesn’t include: (.*)$:捕获doesn’t include:后到文本结尾的所有内容

扩展适配

如果文本中出现其他转折词(如yet),只需修改正则中的转折词列表:(but|however|yet);若存在无排除内容的条目,可添加fillna("")处理空值:

df["column_exclude"] = df["column_exclude"].fillna("")

内容的提问来源于stack exchange,提问作者Paradisum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 09:15:44