如何为数据表中Status列的循环生成计数列N(支持ARRET状态归属到相邻循环)
Got it, let's solve this problem where we need to assign a cycle number N to each row—each cycle starts with FORWARD and ends with REVERSE, and ARRET rows can be linked to either the previous or next cycle. I'll use pandas (the go-to tool for this kind of data manipulation) to walk through the solution step by step.
Step 1: Define the Cycle Logic
- A cycle kicks off whenever
FORWARDappears after a non-FORWARDstatus (likeREVERSEorARRET). - All rows between one
FORWARDstart and the nextFORWARDstart belong to the same cycle (includingREVERSEandARRETrows in between). - We’ll first cover assigning
ARRETrows to the previous cycle (we’ll show the alternative approach later).
Step 2: Code Implementation
First, assume your data is loaded into a pandas DataFrame. If you’re reading from a CSV, use pd.read_csv(); if it’s already in memory, skip that step.
import pandas as pd # Load your data (replace with your actual file path/data source) df = pd.read_csv("your_data_file.csv") # 1. Mark the start of each new cycle # A new cycle starts when 'Status' is 'FORWARD' AND the previous row's status isn't 'FORWARD' # We use fillna('') to handle the first row (no prior row to compare) df["new_cycle"] = (df["Status"] == "FORWARD") & (~df["Status"].shift(1).fillna("").eq("FORWARD")) # 2. Calculate cycle number N by cumulatively summing the new_cycle markers df["N"] = df["new_cycle"].cumsum() # Optional: Drop the temporary 'new_cycle' column if you don't need it df = df.drop(columns=["new_cycle"])
Step 3: How This Works
- The
new_cyclecolumn gets aTrueevery time a fresh cycle starts (aFORWARDafter any other status). - Using
cumsum()on this column increments the cycle number each time a new cycle begins, so all rows between twoFORWARDstarts get the sameNvalue. - For your sample data:
- Rows 484–657 (including the
ARRETblocks at 584–586 and 641–643) getN=1 - Rows 658 onwards (including the final
ARRETrows at 9220–9222) getN=2
- Rows 484–657 (including the
Alternative: Assign ARRET to the Next Cycle
If you want ARRET rows to belong to the upcoming cycle instead, adjust the logic to mark new cycles right after ARRET or REVERSE:
# Mark new cycle when 'Status' is 'FORWARD' and the previous row was 'ARRET'/'REVERSE' df["new_cycle"] = (df["Status"] == "FORWARD") & (df["Status"].shift(1).fillna("").isin(["ARRET", "REVERSE"])) df["N"] = df["new_cycle"].cumsum()
This way, ARRET rows immediately before a new FORWARD will be included in the next cycle rather than the previous one.
Sample Output
After running the first code snippet, your DataFrame will match your expected output:
| Index | X | Z(mm) | FT | Status | N |
|---|---|---|---|---|---|
| 484 | 0.5 | 0.053 | 0.212 | FORWARD | 1 |
| 485 | 0.5 | 0.053 | 0.214 | FORWARD | 1 |
| ... | |||||
| 587 | 0.51 | 0.056 | 0.213 | REVERSE | 1 |
| ... | |||||
| 658 | 0.0 | -0.018 | -0.202 | FORWARD | 2 |
| ... | |||||
| 9222 | 1.01 | -0.14 | 0.117 | ARRET | 2 |
内容的提问来源于stack exchange,提问作者شاه نواز

