使用线性回归填充Pandas数据遇真值错误,求排查解决
Let's break down the issues in your code and fix them step by step:
1. The Ambiguous Truth Value Error
The primary problem here is incorrect indexing in your if condition. When you write data.iloc[i:-1], you're selecting all rows from index i to the second-last row (and all columns), which returns a DataFrame/Series—not a single value. Comparing this to '0' triggers the ValueError because Python can't resolve a single boolean truth value for an entire DataFrame.
You need to target the specific column you want to check for zero values. Let's assume you're imputing values in the last column (adjust this to your actual target column). Replace data.iloc[i:-1] with data.iloc[i, -1] to get the value of the target column in row i.
Also, if your data is numeric, comparing to the string '0' is incorrect—use 0 instead.
2. Assigning to the Wrong Slice
Similarly, when you do data.iloc[i:-1] = lm.predict(x) or data.iloc[i:-1] = [data.iloc[i, -3]/100], you're overwriting multiple rows instead of just the single cell in row i of your target column. Again, use data.iloc[i, target_col_index] (replace with your actual column index) to assign values to the correct cell.
3. Loop Variable "i" Undefined?
This confusion likely stems from the error stopping the loop on the first iteration (i=0). Once we fix the indexing issues, the loop will run correctly, and i will be defined for all iterations.
Corrected Code
Here's the revised code with these fixes (adjust target_col_index to match your actual column index for imputation):
# Define the index of the column you want to impute (e.g., -1 for last column) target_col_index = -1 for i in range(len(data)): # Check if the target column value in row i is 0 (numeric comparison) if data.iloc[i, target_col_index] == 0: # Create input dataframe for prediction x = pd.DataFrame({ 'perception_score': [data.iloc[i, -6]], 'Rating_new': [data.iloc[i, -3]/100], 'Experience': [data.iloc[i, -12]/66] }) # Predict and assign to the target cell (extract scalar with [0]) data.iloc[i, target_col_index] = lm.predict(x)[0] else: # Keep or update the value as intended (matches your original logic) data.iloc[i, target_col_index] = data.iloc[i, -3]/100
Additional Notes
- If your target column uses string values (e.g.,
'0'instead of0), keep the comparison as== '0', but ensure consistency with your data types. lm.predict(x)returns a Series, so adding[0]extracts the scalar value needed to assign to a single cell.- For larger datasets, consider vectorized operations instead of a loop for better performance, but this fixed loop will work reliably for smaller datasets.
内容的提问来源于stack exchange,提问作者Danish Xavier

