如何在Pandas DataFrame中随机选取连续行修改发电量与状态?
Hey there! I totally get why you need consecutive rows for your master's thesis data—your current code picks 12 random scattered rows, which doesn't fit the 2-hour continuous window requirement. Let's fix this step by step:
The Core Issue
df.sample(12) selects 12 random, non-consecutive rows. Instead, we need to pick a single starting point and grab the next 11 rows (since each row is 10 minutes, 12 rows = 2 hours).
Step-by-Step Solution
First, make sure you import random (we'll need it to generate a random starting index):
import random import pandas as pd
Then, calculate the valid range for your starting index. To avoid going out of bounds, the last possible start index is len(df) - 12 (since we need 12 consecutive rows):
# Get the maximum valid starting index (so we don't exceed the DataFrame length) max_start_idx = len(df) - 12 # Generate a random starting index within the valid range start_idx = random.randint(0, max_start_idx)
Now, update the status column and reduce the output (note: you mentioned 'output' in your question but used 'production' in code—adjust to match your actual column name!) for the 12 consecutive rows:
# Set status to 'reduced' for the 12 consecutive rows df.loc[start_idx:start_idx+11, 'status'] = 'reduced' # Reduce output to 60% of its original value df.loc[start_idx:start_idx+11, 'output'] *= 0.6
What if Your DataFrame Uses a Timestamp Index?
If your index is a datetime (since each row represents 10-minute intervals), you can pick a random starting timestamp instead:
# Convert index to a list of timestamps timestamps = df.index.tolist() # Find the latest valid starting index for a 2-hour window max_start_idx = len(timestamps) - 12 # Pick a random starting timestamp start_ts = timestamps[random.randint(0, max_start_idx)] # Calculate the end timestamp (2 hours after the start) end_ts = start_ts + pd.Timedelta(hours=2) # Update the rows in this time window df.loc[start_ts:end_ts, 'status'] = 'reduced' df.loc[start_ts:end_ts, 'output'] *= 0.6
Quick Verification
To confirm you've modified consecutive rows, print the selected range:
print(df.iloc[start_idx:start_idx+12])
This will show you exactly the 12 rows you adjusted, so you can double-check they're continuous.
内容的提问来源于stack exchange,提问作者Elias

