请求提供Fuzzwuzzy包ratio、partial_ratio、token_set_ratio的数学公式
Hey there! Let's dive into the math behind Fuzzwuzzy's ratio, partial_ratio, and token_set_ratio functions, plus walk through a manual calculation for partial_ratio with your example strings: "Mark!" and "mark".
1. ratio函数
This is the most basic function, rooted in the Levenshtein Distance (the minimum number of single-character edits—insertions, deletions, substitutions—needed to turn one string into another).
计算公式:
ratio = (1 - Levenshtein_Distance(str1, str2) / max(len(str1), len(str2))) * 100
- First, calculate the Levenshtein distance between the two strings.
- Divide that distance by the length of the longer string (to normalize against the maximum possible edits needed).
- Subtract the result from 1 to get a similarity fraction, then multiply by 100 to get a percentage score.
2. partial_ratio函数
This function focuses on finding the best matching substring of the shorter string within the longer one, then computes the ratio for that optimal pair. It’s perfect for cases where one string is a subset of the other.
计算步骤:
- Identify the shorter string (
S) and longer string (L) between the two inputs. - Generate all possible substrings of
Lthat have the same length asS(using a sliding window approach). - Calculate the
ratio(using the formula above) for each substring againstS. - The
partial_ratioresult is the maximumratiovalue from all these calculations.
手动计算示例:"Mark!" vs "mark"
Let’s walk through this step by step:
- Fuzzwuzzy defaults to case-insensitive matching, so we first convert both strings to lowercase:
"mark!"(length 5) and"mark"(length 4). - Here,
S = "mark"(shorter, length 4),L = "mark!"(longer, length 5). - Generate all length-4 substrings from
L:- Substring 1:
mark(characters 0–3 of"mark!") - Substring 2:
ark!(characters 1–4 of"mark!")
- Substring 1:
- Calculate
ratiofor each substring vsS:- For
markvsmark: Levenshtein distance is 0. Ratio =(1 - 0/4) * 100 = 100. - For
ark!vsmark: Levenshtein distance is 4 (every character differs). Ratio =(1 - 4/4) * 100 = 0.
- For
- The maximum ratio is 100, so
partial_ratio("Mark!", "mark") = 100.
3. token_set_ratio函数
This function shines for strings with the same core content but different word order or extra tokens. It splits strings into "tokens" (words, split by spaces), uses set operations to isolate overlapping content, then computes ratios on combined token sets.
计算步骤:
- Split both input strings into tokens (default: split on spaces), convert to lowercase, and remove punctuation.
- Compute three token sets:
- Intersection: Tokens present in both strings.
- Difference 1: Tokens only in the first string.
- Difference 2: Tokens only in the second string.
- Generate three combined token strings:
- Intersection + Difference 1
- Intersection + Difference 2
- Intersection alone
- Calculate the
ratiofor every pair of these combined strings. - The
token_set_ratioresult is the maximumratiofrom all these calculations.
核心逻辑:
By prioritizing overlapping tokens, this function ignores redundant or out-of-order words, making it far more robust for messy text like user input or unstructured data.
内容的提问来源于stack exchange,提问作者Neoooar

