You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame二次搜索实现求助:先查Col1再搜Stats列

解决Pandas中先定位特定行再后续搜索的问题

Hey Jeremy, I get why iterrows() was tripping you up—it’s easy to get stuck tracking indices and positions when you’re iterating row by row. Let’s ditch the slow loop approach and use Pandas’ built-in vectorized operations to make this clean and efficient.

核心思路

First, we’ll locate the starting row where Col1 equals "Item327", then search the Stats column after that row for "no" values. We’ll also cover how to loop this logic for multiple matches.

示例代码与分步说明

Let’s start with a sample DataFrame to mirror your data:

import pandas as pd

# 模拟你的数据集
df = pd.DataFrame({
    "Col1": ["Item1", "Item327", "Item5", "Item8", "Item9", "Item10", "Item327", "Item15"],
    "Stats": ["yes", "yes", "yes", "yes", "yes", "no", "yes", "no"]
})

1. 单次搜索(找到第一个"Item327"后的第一个"no")

# 找到第一个"Item327"的索引
start_idx = df[df["Col1"] == "Item327"].index[0]

# 从起始索引的下一行开始筛选Stats为"no"的行
target_row = df.loc[start_idx+1:][df["Stats"] == "no"].iloc[0]

# 添加到新数组(这里用列表存储行数据)
target_array = [target_row]

2. 循环处理所有"Item327"实例

If you have multiple "Item327" entries and need to find the first "no" after each one:

target_array = []

# 获取所有"Item327"的索引列表
item_327_indices = df[df["Col1"] == "Item327"].index.tolist()

for idx in item_327_indices:
    # 从当前索引之后的行中搜索
    post_item_rows = df.loc[idx+1:]
    no_matches = post_item_rows[post_item_rows["Stats"] == "no"]
    
    # 如果找到匹配项,加入目标数组(这里取第一个匹配,要所有就用extend)
    if not no_matches.empty:
        target_array.append(no_matches.iloc[0])

3. 连续搜索"no"(找到一个后继续从下一行找下一个)

If you need to keep searching for "no" values after the first match (instead of restarting at the next "Item327"):

target_array = []
current_start_idx = df[df["Col1"] == "Item327"].index[0]

while current_start_idx < len(df):
    # 从当前位置之后搜索
    subset = df.loc[current_start_idx+1:]
    no_matches = subset[subset["Stats"] == "no"]
    
    if no_matches.empty:
        break  # 没有更多匹配,退出循环
    
    first_no_idx = no_matches.index[0]
    target_array.append(df.loc[first_no_idx])
    current_start_idx = first_no_idx  # 更新起始位置,继续下一轮搜索

为什么不用iterrows()?

  • 效率低: iterrows() is slow for large DataFrames because it iterates row by row instead of using Pandas' optimized vectorized operations.
  • 索引混乱: Tracking positions manually while iterating leads to easy off-by-one errors, which is probably where you got stuck.

最终输出

Your target_array will contain the rows you need (like the 6th row in your example) as Pandas Series objects, which you can convert to dictionaries or lists if needed:

# 转为字典列表方便后续处理
target_dicts = [row.to_dict() for row in target_array]

内容的提问来源于stack exchange,提问作者JeremyV_07

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:18:28