求助:循环中使用.loc/.iloc无法更新Pandas DataFrame值
Fixing DataFrame Update Issues with .iloc in Loops
Hey there! Let's break down why your updates to the 9th column (0-based index 9) aren't sticking, and get this working properly.
The Core Problem with Your Original Code
When you're using iterrows() and modifying the DataFrame mid-loop, you run into two main pitfalls:
- Copy vs. View Confusion: Slicing the DataFrame with
stations[(index+1):min((index+10), len(stations))]usually creates a copy of the original data. Even though you're usingilocto target the original DataFrame, this iterative modification pattern can lead to unexpected behavior because Pandas doesn't guarantee immediate, consistent updates to the underlying data. - Outdated Row Copies:
iterrows()returns copies of each row, not references to the actual data in the DataFrame. So checks likenp.isnan(row.label)might be looking at stale values instead of the updated ones you've already set.
A Reliable Fix: Use an Array to Store Updates
The best practice here is to calculate all your label values first in a separate array, then assign the entire array to your DataFrame column at once. This avoids copy/view headaches and is more efficient.
Here's the revised code:
import numpy as np from shapely.geometry import Polygon # First, make sure we're working with a copy of the DataFrame to avoid view issues stations = stations.copy() # Initialize an array to hold our label values (starts with NaN) labels = np.full(len(stations), np.nan) k = 1 # Iterate over the DataFrame's index instead of using iterrows() for idx in stations.index: # Only process rows that haven't been labeled yet if np.isnan(labels[idx]): # Get the current row's polygon current_row = stations.iloc[idx] p1 = Polygon([ (current_row.long, current_row.lat), (current_row.long1, current_row.lat1), (current_row.long2, current_row.lat2) ]) # Assign the current label to this row labels[idx] = k # Define the range of rows to check (next 9 rows, or end of DataFrame) start_check = idx + 1 end_check = min(idx + 10, len(stations)) # Loop through the target rows for idx1 in stations.index[start_check:end_check]: compare_row = stations.iloc[idx1] p2 = Polygon([ (compare_row.long, compare_row.lat), (compare_row.long1, compare_row.lat1), (compare_row.long2, compare_row.lat2) ]) # If polygons intersect, assign the same label if p1.intersects(p2): labels[idx1] = k k += 1 # Finally, assign the labels array to the 9th column (0-based index 9) stations.iloc[:, 9] = labels
Key Improvements
- Separate Storage: Using a NumPy array to track labels ensures we're modifying a simple, mutable structure instead of dealing with Pandas' copy/view quirks.
- Index-Based Iteration: Looping directly over the DataFrame's index is more reliable than
iterrows(), as we're always referencing the exact position in the original data. - Explicit Copy: Starting with
stations = stations.copy()guarantees we're not working with a view of another DataFrame, which can prevent silent failures when modifying data.
Quick Edge Case Check
- If your original
stationsDataFrame was created from a slice of another DataFrame, it might have been a view instead of a copy—this is a common source of update failures. Thecopy()call fixes this. - Double-check that the 9th column (0-based index 9) exists and uses a mutable dtype (like
floator nullable integer, since we're using NaN values).
内容的提问来源于stack exchange,提问作者AlexModestov
相关产品推荐
相关产品推荐

