You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Scipy从两个变长字符串数组生成指定相似度矩阵?

Efficient Levenshtein Similarity Between Two Word Arrays

To compute the similarity between each word in your unrecognized array (arr2) and every word in your correctly spelled array (arr1) without redundant calculations, you can directly iterate over the pairs of words from the two arrays and compute the Levenshtein ratio for each pair. Here's a step-by-step solution:

Step 1: Install Required Library

First, ensure you have the python-Levenshtein package installed (it provides fast, optimized Levenshtein ratio calculations):

pip install python-Levenshtein

Step 2: Code Implementation

import numpy as np
import pandas as pd
from Levenshtein import ratio

# Define your input arrays
arr1 = np.array(['faucet', 'faucets', 'bath', 'parts', 'bathroom'])
arr2 = np.array(['faucett', 'faucetd', 'bth', 'kichen'])

# Calculate similarity matrix: rows = arr2 words, columns = arr1 words
similarity_matrix = np.array([
    [ratio(word_unrecognized, word_correct) for word_correct in arr1]
    for word_unrecognized in arr2
])

# Convert to pandas DataFrame with meaningful labels
similarity_df = pd.DataFrame(
    similarity_matrix,
    index=arr2,
    columns=arr1
)

# Print the result formatted to 6 decimal places (matching your example)
print(similarity_df.round(6))

Output

Running this code will produce exactly the DataFrame you requested:

faucet   faucets      bath     parts  bathroom
faucett  0.923077  0.857143  0.363636  0.333333  0.266667
faucetd  0.923077  0.857143  0.363636  0.333333  0.266667
bth      0.222222  0.200000  0.857143  0.250000  0.545455
kichen   0.333333  0.307692  0.200000  0.000000  0.142857

Why This Avoids Redundancy

Unlike using scipy.spatial.distance.pdist (which computes all pairwise similarities within a single array, including unnecessary comparisons between words in arr1), this approach directly computes only the pairs you care about: every word in arr2 against every word in arr1. This saves computation time and memory, especially with larger word lists.

Optimization for Large Arrays

If you're working with very large datasets, you can speed things up with parallel processing using libraries like swifter:

import swifter

# Parallelized version for large arrays
similarity_df = pd.DataFrame(arr2, columns=['unrecognized']).swifter.apply(
    lambda row: [ratio(row['unrecognized'], word) for word in arr1],
    axis=1,
    result_type='expand'
)
similarity_df.columns = arr1
similarity_df.index = arr2

内容的提问来源于stack exchange,提问作者Jarad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:07:01