在Pandas DataFrame中查找子串并输出匹配行的下一行
解决方案
步骤1:定位所有包含目标子串的行索引
先找出DataFrame中所有包含指定子串的行,获取它们的索引:
import pandas as pd df = pd.read_excel("sample.xlsx") substring = "Size of file is:" # 生成包含子串的行的布尔掩码 mask = df.apply(lambda row: row.astype(str).str.contains(substring, case=False).any(), axis=1) # 获取匹配行的索引列表 match_indices = df[mask].index
步骤2:筛选有效下一行并提取数据
匹配行的下一行索引为match_indices + 1,但要排除超出DataFrame行数的索引(比如最后一行没有下一行):
# 计算下一行索引,过滤掉超出范围的无效索引 next_row_indices = match_indices + 1 next_row_indices = next_row_indices[next_row_indices < len(df)] # 提取所有目标下一行的数据 result_df = df.loc[next_row_indices]
步骤3:仅保留子串出现次数超50次的结果
先统计匹配行的数量,只有次数超过50次时才输出结果:
if len(match_indices) > 50: print(result_df) # 可选:将结果保存到Excel # result_df.to_excel("target_next_rows.xlsx", index=False) else: print("目标子串出现次数未超过50次")
完整代码
import pandas as pd df = pd.read_excel("sample.xlsx") substring = "Size of file is:" # 定位包含子串的行 mask = df.apply(lambda row: row.astype(str).str.contains(substring, case=False).any(), axis=1) match_indices = df[mask].index # 检查出现次数并输出结果 if len(match_indices) > 50: next_row_indices = match_indices + 1 next_row_indices = next_row_indices[next_row_indices < len(df)] result_df = df.loc[next_row_indices] print(result_df) else: print("目标子串出现次数不足50次,无需输出下一行")
内容的提问来源于stack exchange,提问作者Beast
相关产品推荐
相关产品推荐

