You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas DataFrame提取特殊字符内容并ffill填充后删原行

Pandas提取带冒号文本并向前填充后删除原行的解决方案

你要实现的功能是:从DataFrame的col1列中提取含冒号的文本,用ffill(向前填充)把该值填充到下一个含冒号的行出现前的所有行,最后删除那些包含冒号的行。你之前写的代码没得到预期结果,下面是修正后的实现方案:

原始输入DataFrame

import pandas as pd
df = pd.DataFrame({
    'col1': ['White color :', 'I am not really sure how to do this', 'I am not really sure how to do this', 
             'Black color :', 'I am not ready to solve your issue',
           'I am not ready to solve your issue','I am not ready to solve your issue'],
     })

原代码问题点

  • 提取逻辑没有精准匹配带冒号的行,导致非标题行也提取了无效内容
  • 使用ffill(axis=1)是横向填充,而我们需要的是纵向的向前填充,方向错误

正确实现代码

import pandas as pd
df = pd.DataFrame({
    'col1': ['White color :', 'I am not really sure how to do this', 'I am not really sure how to do this', 
             'Black color :', 'I am not ready to solve your issue',
           'I am not ready to solve your issue','I am not ready to solve your issue'],
     })

# 1. 提取带冒号的行里冒号前的文本,其他行设为NaN
df['new_col'] = df['col1'].str.extract(r'^(.+):$', expand=False).str.strip()
# 2. 对new_col列进行纵向向前填充,把NaN替换成上方最近的有效标题
df['new_col'] = df['new_col'].ffill()
# 3. 过滤掉包含冒号的行,重置索引
df = df[~df['col1'].str.contains(':')].reset_index(drop=True)

print(df)

运行结果(符合预期输出)

col1       new_col
0  I am not really sure how to do this   White color
1  I am not really sure how to do this   White color
2   I am not ready to solve your issue   Black color
3   I am not ready to solve your issue   Black color
4   I am not ready to solve your issue   Black color

内容的提问来源于stack exchange,提问作者s nandan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 02:17:37