如何基于Scipy从两个变长字符串数组生成指定相似度矩阵?
To compute the similarity between each word in your unrecognized array (arr2) and every word in your correctly spelled array (arr1) without redundant calculations, you can directly iterate over the pairs of words from the two arrays and compute the Levenshtein ratio for each pair. Here's a step-by-step solution:
Step 1: Install Required Library
First, ensure you have the python-Levenshtein package installed (it provides fast, optimized Levenshtein ratio calculations):
pip install python-Levenshtein
Step 2: Code Implementation
import numpy as np import pandas as pd from Levenshtein import ratio # Define your input arrays arr1 = np.array(['faucet', 'faucets', 'bath', 'parts', 'bathroom']) arr2 = np.array(['faucett', 'faucetd', 'bth', 'kichen']) # Calculate similarity matrix: rows = arr2 words, columns = arr1 words similarity_matrix = np.array([ [ratio(word_unrecognized, word_correct) for word_correct in arr1] for word_unrecognized in arr2 ]) # Convert to pandas DataFrame with meaningful labels similarity_df = pd.DataFrame( similarity_matrix, index=arr2, columns=arr1 ) # Print the result formatted to 6 decimal places (matching your example) print(similarity_df.round(6))
Output
Running this code will produce exactly the DataFrame you requested:
faucet faucets bath parts bathroom faucett 0.923077 0.857143 0.363636 0.333333 0.266667 faucetd 0.923077 0.857143 0.363636 0.333333 0.266667 bth 0.222222 0.200000 0.857143 0.250000 0.545455 kichen 0.333333 0.307692 0.200000 0.000000 0.142857
Why This Avoids Redundancy
Unlike using scipy.spatial.distance.pdist (which computes all pairwise similarities within a single array, including unnecessary comparisons between words in arr1), this approach directly computes only the pairs you care about: every word in arr2 against every word in arr1. This saves computation time and memory, especially with larger word lists.
Optimization for Large Arrays
If you're working with very large datasets, you can speed things up with parallel processing using libraries like swifter:
import swifter # Parallelized version for large arrays similarity_df = pd.DataFrame(arr2, columns=['unrecognized']).swifter.apply( lambda row: [ratio(row['unrecognized'], word) for word in arr1], axis=1, result_type='expand' ) similarity_df.columns = arr1 similarity_df.index = arr2
内容的提问来源于stack exchange,提问作者Jarad

