如何重置循环?求助:生成随drug变量变化重置的数据集time计数变量
time for drug Groups Hey there! I get exactly what you're trying to do—you need a time variable that counts how many times each drug value has appeared in sequence, and resets the count every time the drug changes. The problem with relying on a row index i is that it doesn't naturally track when the drug group switches, so your counter can't reset properly. Let's go through a few solid solutions depending on the tool you're using.
1. Python with Pandas (Most Efficient)
Pandas has built-in functions to handle this kind of grouped counting without messy loops. We can use groupby combined with cumcount() after creating a unique identifier for each consecutive drug group:
import pandas as pd # Assume your dataset is stored in a DataFrame called df df['time'] = df.groupby((df['drug'] != df['drug'].shift()).cumsum()).cumcount() + 1
How this works:
df['drug'] != df['drug'].shift()creates a boolean series whereTruemeans the current row'sdrugis different from the previous one.- Adding
.cumsum()turns this into a numeric group ID—each time the drug changes, the ID increments. cumcount()then counts the position of each row within its group, starting at 0. Adding 1 makes the count start at 1 (adjust if you want 0-based counting).
2. R with dplyr
If you're working in R, the dplyr package makes this just as straightforward:
library(dplyr) # Assume your dataset is a data frame called df df <- df %>% # Create a temporary group ID for consecutive drug values mutate(group_id = cumsum(drug != lag(drug, default = first(drug)))) %>% # Count rows within each group group_by(group_id) %>% mutate(time = row_number()) %>% ungroup() %>% # Optional: Remove the temporary group_id column select(-group_id)
How this works:
lag(drug, default = first(drug))gets the previous row'sdrugvalue (using the first row's value as the default for the first entry).cumsum(drug != lag(...))creates a group ID that increments every time the drug changes.row_number()counts the position of each row in its group, giving you the resetting counter.
3. Manual Loop (If You Prefer Explicit Control)
If you want to stick with a loop (like your initial attempt), the fix is to track the current drug and a separate counter variable instead of relying on the row index. Here's how to do it in Python:
# Initialize variables to track state current_drug = None counter = 0 time_values = [] # Iterate through each drug value in your dataset for drug in df['drug']: if drug != current_drug: # Reset counter when drug changes current_drug = drug counter = 1 else: # Increment counter if drug stays the same counter += 1 time_values.append(counter) # Assign the list to your DataFrame df['time'] = time_values
This approach avoids the issue with the row index because the counter variable is independent—you can reset it anytime the drug changes, regardless of the row number.
内容的提问来源于stack exchange,提问作者Houssem Missaoui

