如何在Pandas dataframe中查找子字符串的前驱与后继字符串
Pandas查找指定子串对应前驱、后继字符串实现方法
需求对应示例效果如下:
核心思路
- 先通过字符串匹配方法,定位到包含目标子串的行在DataFrame中的位置
- 基于匹配位置做±1的偏移,分别获取前驱(上一行)、后继(下一行)的对应字符串
- 增加边界判断,避免匹配行在第一行/最后一行时出现索引越界错误
代码实现
前置准备
首先构造测试用DataFrame,你可以替换为自己的数据源:
import pandas as pd # 测试用数据,可替换为实际数据 df = pd.DataFrame({ "text_col": [ "2023年行业报告", "2024年季度营收", "2024年中总结", "2024年年度规划", "2025年战略布局" ] }) # 要查找的目标子串 target = "2024年中总结"
场景1:默认连续整数索引
如果你的DataFrame用的是默认生成的0开始的连续整数索引,可以直接用索引值偏移:
# 定位匹配行的索引 match_idx = df[df["text_col"].str.contains(target, na=False)].index[0] # 取值+边界判断 prev_str = df.loc[match_idx - 1, "text_col"] if match_idx > 0 else "无可用前驱" next_str = df.loc[match_idx + 1, "text_col"] if match_idx < len(df) - 1 else "无可用后继" # 输出结果 print(f"匹配字符串:{df.loc[match_idx, 'text_col']}") print(f"前驱字符串:{prev_str}") print(f"后继字符串:{next_str}")
场景2:自定义/非连续索引
如果你的DataFrame用的是自定义索引,或者索引不连续,推荐用位置下标iloc取值,避免索引不连续导致的报错:
# 定位匹配行的位置下标 match_mask = df["text_col"].str.contains(target, na=False) match_pos = df.reset_index().index[match_mask][0] # 取值+边界判断 prev_str = df.iloc[match_pos - 1]["text_col"] if match_pos > 0 else "无可用前驱" next_str = df.iloc[match_pos + 1]["text_col"] if match_pos < len(df) - 1 else "无可用后继"
场景3:存在多个匹配行
如果目标子串匹配到多行,可以遍历所有匹配结果批量输出:
match_masks = df["text_col"].str.contains(target, na=False) match_pos_list = df.reset_index().index[match_masks] for pos in match_pos_list: current = df.iloc[pos]["text_col"] prev = df.iloc[pos-1]["text_col"] if pos > 0 else "无可用前驱" next_ = df.iloc[pos+1]["text_col"] if pos < len(df) -1 else "无可用后继" print(f"当前匹配:{current},前驱:{prev},后继:{next_}")
补充说明
- 如果需要精确匹配完整字符串,把
str.contains(target)替换为== target即可 str.contains方法支持正则表达式匹配,需要匹配特殊字符时可以加regex=False参数关闭正则模式
内容的提问来源于stack exchange,提问作者ML_User
相关产品推荐
相关产品推荐

