n与m个1500维binary vectors集合的相似度度量选型咨询
Great question! Jaccard similarity is actually a perfect fit for your binary vector scenario, and here’s how to adapt it (plus alternatives) to compute the total distance metrics you need to compare set similarity.
Using Jaccard Distance for Total Set Dissimilarity
Since you’re looking for a distance metric (to sum up overall dissimilarity), we’ll use Jaccard Distance—which is just 1 - Jaccard Similarity. Jaccard similarity focuses on the overlap of 1s between two binary vectors, making it ideal when 0s represent the absence of a feature (rather than a meaningful state).
How to compute the total distance for a set:
For a set (e.g., set A with n vectors):
- Calculate the Jaccard Distance between every unique pair of vectors. For n vectors, there are
n*(n-1)/2unique pairs (no need to compare a vector to itself, and comparing i→j is the same as j→i). - Sum all these pairwise distances to get
total_distance_of_n_vectors. - Repeat this process for set B (m vectors) to get
total_distance_of_m_vectors.
As you specified: if total_distance_of_n_vectors > total_distance_of_m_vectors, set B’s vectors are more similar overall (since their total dissimilarity sum is smaller).
Alternative: Hamming Distance
If every bit (0 or 1) represents a meaningful binary attribute (e.g., a sensor reading that’s either on or off), Hamming Distance is a simpler, more intuitive option:
- Hamming Distance counts the number of positions where two binary vectors differ.
- Follow the same aggregation steps: compute all pairwise Hamming Distances, sum them to get the total distance for each set.
- Again, a smaller total distance means higher overall similarity in the set.
Pro Tips for Fair Comparison:
- Normalize for different set sizes: If n ≠ m, divide the total distance by the number of pairs (
k*(k-1)/2for a set of size k) to get an average pairwise distance. This removes bias from sets having more pairs to sum. For example:avg_distance_A = total_distance_of_n_vectors / (n*(n-1)/2) avg_distance_B = total_distance_of_m_vectors / (m*(m-1)/2) - Speed up computations: For large datasets, computing all pairwise distances can be slow. Use vectorized libraries like
scipy.spatial.distance(in Python) to compute pairwise distances efficiently in bulk.
内容的提问来源于stack exchange,提问作者tarun14110

