You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复KNN类中get_ids_of_k_closest的Python名称错误

Fixing the NameError and Other Issues in Your KNN Implementation

Hey there, let's tackle that NameError first—it's a common gotcha with Python class methods, and then we'll iron out a few other bugs in your code to make your KNN work properly.

Core Fix for the NameError

The get_ids_of_k_closest method is an instance method of your simple_KNN class, but you're treating it like a standalone function. Here's what's wrong and how to fix it:

  • When defining instance methods in a class, the first parameter must always be self (this refers to the class instance itself). Your current definition misses self.
  • When calling the method from another instance method (like predict_new_item), you need to prefix it with self. to access it from the instance.
  • Also, you're using a variable k directly, but your class stores k as an instance attribute (self.k), so you should use that instead of an undefined global k.

Corrected Method Definition & Call

# Inside the simple_KNN class, fix get_ids_of_k_closest definition
def get_ids_of_k_closest(self, distFromNewItem, k):
    closestK = np.empty(k, dtype=int)
    arraySize = distFromNewItem.shape[0]  # Fixed array size calculation
    for i in range(k):  # Renamed loop variable to avoid overriding parameter k
        thisClosest = 0
        for exemplar in range(arraySize):
            if distFromNewItem[exemplar] < distFromNewItem[thisClosest]:
                thisClosest = exemplar
        closestK[i] = thisClosest
        distFromNewItem[thisClosest] = 100000
    return closestK

# Inside predict_new_item, fix the call
closestK = self.get_ids_of_k_closest(distFromNewItem, self.k)

Other Critical Bugs to Fix

Your code has a few more issues that will prevent it from working even after fixing the NameError:

  1. Loop Range Errors

    • All instances of range(x-1) will skip the last element in your arrays. For example, for stored_example in range(self.numTrainingItems-1) misses the last training sample. Change these to range(self.numTrainingItems) (and similar for other loops).
  2. Incorrect Attribute Name

    • You're using self.y_train in predict_new_item, but in your fit method you assigned the labels to self.modelY. Replace self.y_train with self.modelY.
  3. Label Counting Logic

    • labelCounts[thisindex] +=1 is wrong: thisindex is the index of the training sample, not the index of the label in self.labelsPresent. You need to map the sample's label to its position in self.labelsPresent:
      this_label = self.modelY[thisindex]
      label_idx = np.where(self.labelsPresent == this_label)[0][0]
      labelCounts[label_idx] += 1
      
  4. Prediction Calculation

    • np.amax(labelCounts) returns the highest count, but you need the corresponding label. Use np.argmax to get the index of the highest count, then fetch the label from self.labelsPresent:
      thisPrediction = self.labelsPresent[np.argmax(labelCounts)]
      
  5. Feature Count Bug in fit

    • self.numFeatures = X.shape[0] is incorrect—X.shape[0] is the number of samples. You want X.shape[1] for the number of features.
  6. Variable Name Conflict

    • In get_ids_of_k_closest, your loop variable k overrides the method's k parameter. Rename the loop variable to something like i to avoid this.

Full Corrected Code

import numpy as np
# Assuming W7utils is available with euclidean_distance
# import W7utils

class simple_KNN:
    def __init__(self, k, verbose=True):
        self.k = k
        # Replace with your actual distance function if needed
        self.distance = lambda a, b: np.sqrt(np.sum((a - b)**2))  # Example euclidean distance
        self.verbose = verbose

    def fit(self, X, y):
        self.numTrainingItems = X.shape[0]
        self.numFeatures = X.shape[1]  # Fixed feature count
        self.modelX = X
        self.modelY = y
        self.labelsPresent = np.unique(self.modelY)
        if self.verbose:
            print( f"共有{self.numTrainingItems}个训练样本,每个样本由{self.numFeatures}个特征值描述")
            print( f"因此self.modelX是一个形状为{self.modelX.shape}的二维数组")
            print( f"self.modelY是一个包含{len(self.modelY)}个元素的列表,每个元素是以下标签之一{self.labelsPresent}")

    def predict(self, newItems):
        numToPredict = newItems.shape[0]
        predictions = np.empty(numToPredict)
        for item in range(numToPredict):  # Fixed loop range
            thisPrediction = self.predict_new_item(newItems[item])
            predictions[item] = thisPrediction
        return predictions

    def predict_new_item(self, newItem):
        distFromNewItem = np.zeros((self.numTrainingItems))
        for stored_example in range(self.numTrainingItems):  # Fixed loop range
            distFromNewItem[stored_example] = self.distance(newItem, self.modelX[stored_example])
        # Call the method with self and use instance's k
        closestK = self.get_ids_of_k_closest(distFromNewItem, self.k)
        labelCounts = np.zeros(len(self.labelsPresent))
        for i in range(self.k):  # Fixed loop range and variable name
            thisindex = closestK[i]
            this_label = self.modelY[thisindex]  # Fixed attribute name
            # Map label to its index in labelsPresent
            label_idx = np.where(self.labelsPresent == this_label)[0][0]
            labelCounts[label_idx] += 1  # Fixed counting logic
        # Get the label with highest count
        thisPrediction = self.labelsPresent[np.argmax(labelCounts)]
        return thisPrediction

    def get_ids_of_k_closest(self, distFromNewItem, k):  # Added self parameter
        closestK = np.empty(k, dtype=int)
        arraySize = distFromNewItem.shape[0]  # Fixed array size
        for i in range(k):  # Renamed loop variable to avoid conflict
            thisClosest = 0
            for exemplar in range(arraySize):  # Fixed loop range
                if distFromNewItem[exemplar] < distFromNewItem[thisClosest]:
                    thisClosest = exemplar
            closestK[i] = thisClosest
            distFromNewItem[thisClosest] = 100000
        return closestK

Notes

  • I replaced W7utils.euclidean_distance with a lambda function for testing—swap it back to your actual utility function if needed.
  • The core fix for the NameError was adding self to the method definition and using self.get_ids_of_k_closest() when calling it.

内容的提问来源于stack exchange,提问作者William Barnes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:38:14