求助:编写支持保序部分匹配且避免过度匹配的通用正则表达式生成方案
Let's break down your problem and fix this step by step. Your core need is to find ordered partial matches between a field's keywords and a target text—meaning text keywords must appear in the same order as they do in the field, but you don't need to match every keyword, and you want to avoid false matches with unrelated content.
First, let's identify the flaws in your original code:
- When you concatenate all split keywords without optional markers (
?), the regex requires every keyword from the field to exist in the text, which breaks partial matching (like your first test case would fail). - Adding
?to make each keyword optional creates a regex that matches any text (since all groups are optional, even empty strings pass), leading to false positives like your third test case.
Solution 1: Subsequence Check (Recommended)
Instead of wrestling with regex edge cases, a simpler and more reliable approach is to check if the text's keywords form an ordered subsequence of the field's keywords. This is straightforward to implement and easy to debug.
Code Implementation
def match_ordered_partial(field, text): import re # Helper to extract lowercase, non-empty keywords from a string def extract_keywords(s): return [k.lower() for k in re.split(r'\W+', s) if k] field_keywords = extract_keywords(field) text_keywords = extract_keywords(text) # If text has no valid keywords, immediately return False if not text_keywords: return False # Check if text keywords are an ordered subsequence of field keywords field_iter = iter(field_keywords) return all(key in field_iter for key in text_keywords) # Test your examples print(match_ordered_partial("F1 gearbox: 0-400 m","f1 gearbox")) # ✅ True (partial match, correct order) print(match_ordered_partial("0-100 kmph" , "100-0 kmph")) # ❌ False (wrong order) print(match_ordered_partial("F1 gearbox: 0-400 m","none")) # ❌ False (no matching keywords) print(match_ordered_partial("Combined* (ECE+EUDC) (l/100 km)","combined ece eudc")) # ✅ True (partial match, ignores special chars)
How It Works
- Keyword Extraction: We split both inputs using non-word characters (
\W+), filter out empty strings, and convert to lowercase for case-insensitive matching. - Subsequence Validation: Using an iterator over the field's keywords, we check if each text keyword appears in the field in order. The iterator ensures we only look forward in the field's keyword list, preserving the required sequence.
Solution 2: Regex-Based Approach
If you specifically need a regex solution, we can generate a pattern that enforces ordered partial matches while avoiding false positives. The key is to ensure at least one keyword is matched, and any additional keywords follow the field's order.
Code Implementation
def regexp(field, text): import re # Extract and escape keywords from field (handles special regex chars) field_keywords = [re.escape(k.lower()) for k in re.split(r'\W+', field) if k] if not field_keywords: return False # Build regex parts: each part matches a starting keyword + optional subsequent keywords (in order) regex_parts = [] for start_idx in range(len(field_keywords)): current_part = field_keywords[start_idx] # Add optional matches for keywords that come after the starting one for next_idx in range(start_idx + 1, len(field_keywords)): current_part += f"(.*{field_keywords[next_idx]})?" regex_parts.append(current_part) # Combine parts with OR, wrap to allow any characters around matches full_regex = fr".*({'|'.join(regex_parts)}).*" pattern = re.compile(full_regex, re.IGNORECASE) match_result = pattern.search(text) print(match_result, "\n", pattern) return bool(match_result) # Test your examples print(regexp("F1 gearbox: 0-400 m","f1 gearbox")) # ✅ True print(regexp("0-100 kmph" , "100-0 kmph")) # ❌ False print(regexp("F1 gearbox: 0-400 m","none")) # ❌ False print(regexp("Combined* (ECE+EUDC) (l/100 km)","combined ece eudc")) # ✅ True
How It Works
- Keyword Escaping:
re.escape()ensures special characters (like*,+, or()) from the field don't break the regex. - Regex Construction: We build parts for every possible starting keyword in the field, each allowing optional matches for subsequent keywords (in their original order). For example, if field keywords are
["f1", "gearbox", "0"], the regex parts are:f1(.*gearbox)?(.*0)?(matches "f1", "f1 gearbox", "f1 ... gearbox ... 0", etc.)gearbox(.*0)?(matches "gearbox", "gearbox 0", etc.)0(matches just "0")
- Final Pattern: Joining these parts with
|(OR) ensures we match any valid ordered partial subset of the field's keywords, and wrapping with.*allows characters before/after the match.
内容的提问来源于stack exchange,提问作者inkarar

