You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何正确使用GitHub上的manhattan_distance函数?调用遇类型错误

Fixing the TypeError in manhattan_distance for Text Strings

Hey there! Let's break down why you're seeing that TypeError and how to correctly use the manhattan_distance function with text.

Why the Error Happens

The manhattan_distance function you're using is built to work with numeric vectors—think arrays of numbers, not strings. When you pass in string arrays or lists of strings, NumPy can't perform subtraction (p_vec - q_vec) on string values, which is exactly what that error is telling you: there's no valid way to subtract strings in the way the function expects.

Correct Usage Steps

To use this function with text, you first need to convert your text into numeric feature vectors. Two common, easy-to-implement methods for this are:

  • Bag of Words: Counts how often each word appears in your text
  • TF-IDF: Measures how important a word is to a document relative to a larger set of text

Let's walk through a complete example using scikit-learn's CountVectorizer (you'll need to install scikit-learn first if you haven't: pip install scikit-learn):

import numpy as np
from sklearn.feature_extraction.text import CountVectorizer

# Recreate the Similarity class from your module
class Similarity:
    def __init__(self, e):
        self.e = e
    
    def manhattan_distance(self, p_vec, q_vec):
        return max(np.sum(np.fabs(p_vec - q_vec)), self.e)

# Step 1: Prepare your text data
text1 = "this is a test"
text2 = "this is a testee"

# Step 2: Convert text to numeric vectors
vectorizer = CountVectorizer()
# Fit the vectorizer to your text and transform it into vectors
text_vectors = vectorizer.fit_transform([text1, text2])
# Convert from sparse matrix to dense NumPy array (required for the distance function)
p_vec = text_vectors[0].toarray().flatten()
q_vec = text_vectors[1].toarray().flatten()

# Step 3: Use the manhattan_distance function
t = Similarity(1)
distance = t.manhattan_distance(p_vec, q_vec)
print(distance) # Output will be 2, since the two vectors differ in two positions

What This Does

  1. The CountVectorizer creates a vocabulary from your text (in this case, the words this, is, a, test, testee).
  2. It converts each sentence into a vector where each position represents the count of a word from the vocabulary:
    • text1 becomes [1, 1, 1, 1, 0] (counts for each word in order)
    • text2 becomes [1, 1, 1, 0, 1]
  3. Now these are valid numeric vectors, so the manhattan_distance function can subtract them, take absolute values, sum the differences, and return the result.

Why Your Previous Calls Failed

  • Your first call passed single-element arrays with full sentences: np.array(["this is a testee"]) is a string array, not a numeric vector the function can process.
  • Your second call passed lists of strings: ["this", "is", "a", "test"] is a list of string values, which can't be subtracted like numbers.

内容的提问来源于stack exchange,提问作者João Pedro Hannauer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:59:18