如何用Pandas DataFrame提取特殊字符内容并ffill填充后删原行
Pandas提取带冒号文本并向前填充后删除原行的解决方案
你要实现的功能是:从DataFrame的col1列中提取含冒号的文本,用ffill(向前填充)把该值填充到下一个含冒号的行出现前的所有行,最后删除那些包含冒号的行。你之前写的代码没得到预期结果,下面是修正后的实现方案:
原始输入DataFrame
import pandas as pd df = pd.DataFrame({ 'col1': ['White color :', 'I am not really sure how to do this', 'I am not really sure how to do this', 'Black color :', 'I am not ready to solve your issue', 'I am not ready to solve your issue','I am not ready to solve your issue'], })
原代码问题点
- 提取逻辑没有精准匹配带冒号的行,导致非标题行也提取了无效内容
- 使用
ffill(axis=1)是横向填充,而我们需要的是纵向的向前填充,方向错误
正确实现代码
import pandas as pd df = pd.DataFrame({ 'col1': ['White color :', 'I am not really sure how to do this', 'I am not really sure how to do this', 'Black color :', 'I am not ready to solve your issue', 'I am not ready to solve your issue','I am not ready to solve your issue'], }) # 1. 提取带冒号的行里冒号前的文本,其他行设为NaN df['new_col'] = df['col1'].str.extract(r'^(.+):$', expand=False).str.strip() # 2. 对new_col列进行纵向向前填充,把NaN替换成上方最近的有效标题 df['new_col'] = df['new_col'].ffill() # 3. 过滤掉包含冒号的行,重置索引 df = df[~df['col1'].str.contains(':')].reset_index(drop=True) print(df)
运行结果(符合预期输出)
col1 new_col 0 I am not really sure how to do this White color 1 I am not really sure how to do this White color 2 I am not ready to solve your issue Black color 3 I am not ready to solve your issue Black color 4 I am not ready to solve your issue Black color
内容的提问来源于stack exchange,提问作者s nandan
相关产品推荐
相关产品推荐

