如何正确使用GitHub上的manhattan_distance函数?调用遇类型错误
manhattan_distance for Text Strings Hey there! Let's break down why you're seeing that TypeError and how to correctly use the manhattan_distance function with text.
Why the Error Happens
The manhattan_distance function you're using is built to work with numeric vectors—think arrays of numbers, not strings. When you pass in string arrays or lists of strings, NumPy can't perform subtraction (p_vec - q_vec) on string values, which is exactly what that error is telling you: there's no valid way to subtract strings in the way the function expects.
Correct Usage Steps
To use this function with text, you first need to convert your text into numeric feature vectors. Two common, easy-to-implement methods for this are:
- Bag of Words: Counts how often each word appears in your text
- TF-IDF: Measures how important a word is to a document relative to a larger set of text
Let's walk through a complete example using scikit-learn's CountVectorizer (you'll need to install scikit-learn first if you haven't: pip install scikit-learn):
import numpy as np from sklearn.feature_extraction.text import CountVectorizer # Recreate the Similarity class from your module class Similarity: def __init__(self, e): self.e = e def manhattan_distance(self, p_vec, q_vec): return max(np.sum(np.fabs(p_vec - q_vec)), self.e) # Step 1: Prepare your text data text1 = "this is a test" text2 = "this is a testee" # Step 2: Convert text to numeric vectors vectorizer = CountVectorizer() # Fit the vectorizer to your text and transform it into vectors text_vectors = vectorizer.fit_transform([text1, text2]) # Convert from sparse matrix to dense NumPy array (required for the distance function) p_vec = text_vectors[0].toarray().flatten() q_vec = text_vectors[1].toarray().flatten() # Step 3: Use the manhattan_distance function t = Similarity(1) distance = t.manhattan_distance(p_vec, q_vec) print(distance) # Output will be 2, since the two vectors differ in two positions
What This Does
- The
CountVectorizercreates a vocabulary from your text (in this case, the wordsthis,is,a,test,testee). - It converts each sentence into a vector where each position represents the count of a word from the vocabulary:
text1becomes[1, 1, 1, 1, 0](counts for each word in order)text2becomes[1, 1, 1, 0, 1]
- Now these are valid numeric vectors, so the
manhattan_distancefunction can subtract them, take absolute values, sum the differences, and return the result.
Why Your Previous Calls Failed
- Your first call passed single-element arrays with full sentences:
np.array(["this is a testee"])is a string array, not a numeric vector the function can process. - Your second call passed lists of strings:
["this", "is", "a", "test"]is a list of string values, which can't be subtracted like numbers.
内容的提问来源于stack exchange,提问作者João Pedro Hannauer

