如何修复KNN类中get_ids_of_k_closest的Python名称错误
Fixing the NameError and Other Issues in Your KNN Implementation
Hey there, let's tackle that NameError first—it's a common gotcha with Python class methods, and then we'll iron out a few other bugs in your code to make your KNN work properly.
Core Fix for the NameError
The get_ids_of_k_closest method is an instance method of your simple_KNN class, but you're treating it like a standalone function. Here's what's wrong and how to fix it:
- When defining instance methods in a class, the first parameter must always be
self(this refers to the class instance itself). Your current definition missesself. - When calling the method from another instance method (like
predict_new_item), you need to prefix it withself.to access it from the instance. - Also, you're using a variable
kdirectly, but your class storeskas an instance attribute (self.k), so you should use that instead of an undefined globalk.
Corrected Method Definition & Call
# Inside the simple_KNN class, fix get_ids_of_k_closest definition def get_ids_of_k_closest(self, distFromNewItem, k): closestK = np.empty(k, dtype=int) arraySize = distFromNewItem.shape[0] # Fixed array size calculation for i in range(k): # Renamed loop variable to avoid overriding parameter k thisClosest = 0 for exemplar in range(arraySize): if distFromNewItem[exemplar] < distFromNewItem[thisClosest]: thisClosest = exemplar closestK[i] = thisClosest distFromNewItem[thisClosest] = 100000 return closestK # Inside predict_new_item, fix the call closestK = self.get_ids_of_k_closest(distFromNewItem, self.k)
Other Critical Bugs to Fix
Your code has a few more issues that will prevent it from working even after fixing the NameError:
Loop Range Errors
- All instances of
range(x-1)will skip the last element in your arrays. For example,for stored_example in range(self.numTrainingItems-1)misses the last training sample. Change these torange(self.numTrainingItems)(and similar for other loops).
- All instances of
Incorrect Attribute Name
- You're using
self.y_traininpredict_new_item, but in yourfitmethod you assigned the labels toself.modelY. Replaceself.y_trainwithself.modelY.
- You're using
Label Counting Logic
labelCounts[thisindex] +=1is wrong:thisindexis the index of the training sample, not the index of the label inself.labelsPresent. You need to map the sample's label to its position inself.labelsPresent:this_label = self.modelY[thisindex] label_idx = np.where(self.labelsPresent == this_label)[0][0] labelCounts[label_idx] += 1
Prediction Calculation
np.amax(labelCounts)returns the highest count, but you need the corresponding label. Usenp.argmaxto get the index of the highest count, then fetch the label fromself.labelsPresent:thisPrediction = self.labelsPresent[np.argmax(labelCounts)]
Feature Count Bug in
fitself.numFeatures = X.shape[0]is incorrect—X.shape[0]is the number of samples. You wantX.shape[1]for the number of features.
Variable Name Conflict
- In
get_ids_of_k_closest, your loop variablekoverrides the method'skparameter. Rename the loop variable to something likeito avoid this.
- In
Full Corrected Code
import numpy as np # Assuming W7utils is available with euclidean_distance # import W7utils class simple_KNN: def __init__(self, k, verbose=True): self.k = k # Replace with your actual distance function if needed self.distance = lambda a, b: np.sqrt(np.sum((a - b)**2)) # Example euclidean distance self.verbose = verbose def fit(self, X, y): self.numTrainingItems = X.shape[0] self.numFeatures = X.shape[1] # Fixed feature count self.modelX = X self.modelY = y self.labelsPresent = np.unique(self.modelY) if self.verbose: print( f"共有{self.numTrainingItems}个训练样本,每个样本由{self.numFeatures}个特征值描述") print( f"因此self.modelX是一个形状为{self.modelX.shape}的二维数组") print( f"self.modelY是一个包含{len(self.modelY)}个元素的列表,每个元素是以下标签之一{self.labelsPresent}") def predict(self, newItems): numToPredict = newItems.shape[0] predictions = np.empty(numToPredict) for item in range(numToPredict): # Fixed loop range thisPrediction = self.predict_new_item(newItems[item]) predictions[item] = thisPrediction return predictions def predict_new_item(self, newItem): distFromNewItem = np.zeros((self.numTrainingItems)) for stored_example in range(self.numTrainingItems): # Fixed loop range distFromNewItem[stored_example] = self.distance(newItem, self.modelX[stored_example]) # Call the method with self and use instance's k closestK = self.get_ids_of_k_closest(distFromNewItem, self.k) labelCounts = np.zeros(len(self.labelsPresent)) for i in range(self.k): # Fixed loop range and variable name thisindex = closestK[i] this_label = self.modelY[thisindex] # Fixed attribute name # Map label to its index in labelsPresent label_idx = np.where(self.labelsPresent == this_label)[0][0] labelCounts[label_idx] += 1 # Fixed counting logic # Get the label with highest count thisPrediction = self.labelsPresent[np.argmax(labelCounts)] return thisPrediction def get_ids_of_k_closest(self, distFromNewItem, k): # Added self parameter closestK = np.empty(k, dtype=int) arraySize = distFromNewItem.shape[0] # Fixed array size for i in range(k): # Renamed loop variable to avoid conflict thisClosest = 0 for exemplar in range(arraySize): # Fixed loop range if distFromNewItem[exemplar] < distFromNewItem[thisClosest]: thisClosest = exemplar closestK[i] = thisClosest distFromNewItem[thisClosest] = 100000 return closestK
Notes
- I replaced
W7utils.euclidean_distancewith a lambda function for testing—swap it back to your actual utility function if needed. - The core fix for the
NameErrorwas addingselfto the method definition and usingself.get_ids_of_k_closest()when calling it.
内容的提问来源于stack exchange,提问作者William Barnes
相关产品推荐
相关产品推荐

