Pandas遍历DataFrame行:实现值右移且不溢出列范围的需求
Looks like you need to populate the subsequent rows of your DataFrame using values from the first row (2018-03-16), with two key rules: every row's day1 value must match the first row's day1, and you can't add new columns (stick to day1 to day5). Let's walk through two common interpretations of your requirement with code examples.
Scenario 1: Replicate the First Row Entirely to All Subsequent Rows
If you want every row after the first to be an exact copy of the first row (so day1 stays consistent, and all other columns use the first row's values), this is straightforward:
Code Implementation
import pandas as pd import numpy as np # Build your initial DataFrame data = { 'day1': [1.0, np.nan, np.nan, np.nan, np.nan], 'day2': [2.0, np.nan, np.nan, np.nan, np.nan], 'day3': [3.0, np.nan, np.nan, np.nan, np.nan], 'day4': [4.0, np.nan, np.nan, np.nan, np.nan], 'day5': [5.0, np.nan, np.nan, np.nan, np.nan] } index = pd.date_range('2018-03-16', periods=5) df = pd.DataFrame(data, index=index) # Extract the first row's values first_row = df.iloc[0] # Assign the first row to all subsequent rows df.loc[df.index[1:]] = first_row.values # Print the result print(df)
Output
day1 day2 day3 day4 day5 2018-03-16 1.0 2.0 3.0 4.0 5.0 2018-03-17 1.0 2.0 3.0 4.0 5.0 2018-03-18 1.0 2.0 3.0 4.0 5.0 2018-03-19 1.0 2.0 3.0 4.0 5.0 2018-03-20 1.0 2.0 3.0 4.0 5.0
Scenario 2: Shift First Row Values Right While Keeping day1 Fixed
If your goal is to "shift" the first row's values right for each subsequent row (with day1 always staying as the first row's day1 value), here's how to do it without exceeding column limits:
Code Implementation
import pandas as pd import numpy as np # Build your initial DataFrame (same as above) data = { 'day1': [1.0, np.nan, np.nan, np.nan, np.nan], 'day2': [2.0, np.nan, np.nan, np.nan, np.nan], 'day3': [3.0, np.nan, np.nan, np.nan, np.nan], 'day4': [4.0, np.nan, np.nan, np.nan, np.nan], 'day5': [5.0, np.nan, np.nan, np.nan, np.nan] } index = pd.date_range('2018-03-16', periods=5) df = pd.DataFrame(data, index=index) # Extract the first row's values as a numpy array first_row_vals = df.iloc[0].values # Get the number of columns to avoid exceeding limits num_cols = df.shape[1] # Iterate over each row after the first for row_idx in range(1, len(df)): # Calculate how many positions to shift (matches the row index, starting at 1) shift_count = row_idx # Build the row values: start with fixed day1, then fill with shifted values # First, add the fixed day1 value new_row = [first_row_vals[0]] # Add the shifted part: repeat day1 shift_count times, then take remaining first row values shifted_part = [first_row_vals[0]] * shift_count # Take the remaining values from the first row (without going over column limits) remaining_vals = first_row_vals[:num_cols - 1 - shift_count] # Combine parts and ensure we don't exceed column count new_row.extend(shifted_part + list(remaining_vals)) # If we're short on values (for later rows), fill with NaN (though in this case, we won't need it) while len(new_row) < num_cols: new_row.append(np.nan) # Assign the new row to the DataFrame df.iloc[row_idx] = new_row # Print the result print(df)
Output
day1 day2 day3 day4 day5 2018-03-16 1.0 2.0 3.0 4.0 5.0 2018-03-17 1.0 1.0 2.0 3.0 4.0 2018-03-18 1.0 1.0 1.0 2.0 3.0 2018-03-19 1.0 1.0 1.0 1.0 2.0 2018-03-20 1.0 1.0 1.0 1.0 1.0
Both solutions respect your requirement to stay within existing column bounds. Pick the one that aligns with your actual desired output!
内容的提问来源于stack exchange,提问作者denister

