You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计数据集中当前行标签为Tag1且前一行标签为Tag2的行数

Fixing Your Code and Efficient Alternatives

First, let's break down why you're hitting that error:
When using iterrows(), row is a pandas Series representing a single row of your DataFrame. Trying to access row[index+1] is asking for a column in that row named after the integer index+1—which doesn't exist (your columns are things like 'tag'), hence the "string indices must be integers" error. You need to fetch the previous row directly from the DataFrame instead of trying to get it from the current row.

Also, note your original code checks the next row instead of the previous one (using index+1), which doesn't match your goal of counting rows where the current tag is "Tag1" and the prior row is "Tag2". Let's fix both issues:

Fixed Loop Code

counter = 0
for index, row in train.iterrows():
    # Skip the first row—it has no previous row to compare
    if index == 0:
        continue
    # Grab the previous row from the original DataFrame
    previous_row = train.iloc[index - 1]
    if row['tag'] == "Tag1" and previous_row['tag'] == "Tag2":
        counter += 1

More Efficient Approach Using shift()

Looping with iterrows() is slow for large datasets. A far better way is to use pandas' vectorized shift() method, which lets you compare rows without manual looping:

# Create a column holding the tag from the row above
train['previous_tag'] = train['tag'].shift(1)

# Count rows where both conditions are true
count = len(train[(train['tag'] == "Tag1") & (train['previous_tag'] == "Tag2")])

If you don't want to add a new column to your DataFrame, you can do it in a single line:

count = ((train['tag'] == "Tag1") & (train['tag'].shift(1) == "Tag2")).sum()

How This Works:

  • train['tag'].shift(1) shifts the 'tag' column down by one row, so each entry now holds the value from the row immediately above it.
  • We create a boolean mask where both your conditions are satisfied, then sum the True values (since True equals 1 in numeric contexts) to get the total count.

This method is way faster than looping, especially for large datasets, because it leverages pandas' optimized internal operations.

内容的提问来源于stack exchange,提问作者Ehsan Mehralian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:39:17