You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas DataFrame中查找特定字符串并拆分生成新行?

解决方案

可以通过正则拆分+行展开的方式实现需求,具体步骤如下:

  1. 导入依赖并定义原始数据
  2. 指定需要作为分隔点的关键词
  3. 构建正则表达式,确保拆分时保留关键词本身
  4. 拆分字符串并展开为多行,同时清理多余空格

完整代码

import pandas as pd
import re

# 原始DataFrame
data = {'Name': ['112 13 15 195 46 60 7 28 88 73 16 91 9  Number of days 32 30 29 32 30 30 31 32 31 29 30 31 30  Appx. average wt 39 4 40 74 331 302 273 263 277 274 307 295 303']}
df = pd.DataFrame(data)

# 定义分隔关键词
split_keywords = ["Number of days", "Appx. average wt"]

# 构建正则匹配规则:匹配关键词前的空格,且保留关键词
split_pattern = r'\s+(?=' + '|'.join(map(re.escape, split_keywords)) + r')'

# 执行拆分、展开、清理操作
result_df = df['Name'].str.split(split_pattern, expand=False).explode().str.strip().to_frame('Name')

# 查看结果
print(result_df)

代码说明

  • re.escape(split_keywords):对关键词中的特殊字符(比如Appx. average wt里的点号)进行转义,避免正则解析错误
  • 正则\s+(?=...):使用正向预查匹配关键词前的连续空格,确保拆分后每个片段都包含对应的关键词
  • explode():将拆分得到的列表元素逐个展开为独立行
  • str.strip():去除每个字符串首尾的多余空格,保证格式整洁

运行后得到的result_df就是你需要的目标DataFrame。

内容的提问来源于stack exchange,提问作者san1

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 01:20:37