You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame中通过字符串匹配将字符串拆分为两个子串

解决方法

可以利用 pandas 的字符串拆分方法 str.split() 实现需求,以下是具体实现步骤:

示例代码

import pandas as pd

# 构造符合逻辑的原数据集 df1(修正原示例中英文字词与中文语句不匹配的问题)
data = {
    "Sentence": ["我和约翰前往该区域,耗时20分钟", "我跑出房子后仓促下了结论"],
    "word": ["前往", "仓促"]
}
df1 = pd.DataFrame(data)

# 按每行对应的word拆分Sentence,生成source和target列
df2 = df1.join(
    df1["Sentence"].str.split(df1["word"], n=1, expand=True).rename(columns={0: "source", 1: "target"})
)

print(df2)

代码说明

  • str.split(df1["word"], n=1, expand=True):以每行专属的word为分隔符,对Sentence执行仅1次拆分,并将拆分结果展开为独立DataFrame
  • rename():将拆分生成的两列重命名为需求指定的source和target
  • join():把拆分后的列合并回原数据集,得到最终的df2

输出结果

Sentencewordsourcetarget
我和约翰前往该区域,耗时20分钟前往我和约翰该区域,耗时20分钟
我跑出房子后仓促下了结论仓促我跑出房子后下了结论

注意:使用时需确保word列的内容确实存在于对应行的Sentence中,否则拆分操作会返回空值。

内容的提问来源于stack exchange,提问作者Vivek kanna Jayaprakash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:41:06