Pandas初学者咨询:iloc列索引、多列选取及回归建模问题
Hey there! Let's break down your questions one by one to help you get a better grip on Pandas and linear regression with scikit-learn:
iloc[:,1] used to locate columns by index? Yes, exactly! iloc[:,1] targets columns using positional indexing. The : means "select all rows", and the 1 refers to the column at position 1 (remember: Pandas uses 0-based indexing, so this is actually the second column in your DataFrame).
You don't need to use iloc[] one by one—slice syntax lets you batch-select columns directly:
# Select columns from position 1 to 10 (inclusive, since iloc slices are left-closed, right-open) selected_columns = your_dataframe.iloc[:, 1:11]
Here's why this works: iloc follows the start:end rule where start is included, but end is not. To include the column at position 10, we set end to 11.
iteritems() error) The error "Series' objects are mutable, thus they cannot be hashed" happens because iteritems() returns column names paired with individual Series (1-dimensional data), but scikit-learn's LinearRegression.fit() expects a 2D feature matrix (like a DataFrame or 2D numpy array), not single Series objects.
You don't need to iterate over columns at all—just pass the subset of columns 0-9 directly as your feature set. Here's how to fix your code:
from sklearn import linear_model import pandas as pd import matplotlib.pyplot as plt regrmodel = linear_model.LinearRegression() print("Y train", y_train) # Select columns 0-9 from X_train as features X_train_features = X_train.iloc[:, 0:10] regrmodel.fit(X_train_features, y_train) # Use the same feature subset for prediction X_test_features = X_test.iloc[:, 0:10] y_test_pred = regrmodel.predict(X_test_features) # Format prediction results and align index y_test_pred = pd.Series(y_test_pred) y_test_pred.index = y_test.index # Visualize (using column 9 as an example) plt.scatter(X_test.iloc[:,9], y_test, color='red', label='Actual data') plt.scatter(X_test.iloc[:,9], y_test_pred, color='green', label='Predicted data') plt.legend() plt.show()
This approach passes a 2D feature matrix directly to the model, which fits its input requirements and avoids the hashing error entirely.
内容的提问来源于stack exchange,提问作者Unknown

