如何在Python中实现类Flex的正则匹配,修正指令/ID/运算符识别问题
Solution
Your original code had three main issues:
re.search()only returns the first match, so you missed all subsequent matches for instructions, operators, and IDs.- The ID regex didn't exclude instruction keywords, leading to false matches on commands like
Set. - The instruction regex ordered shorter patterns first (e.g.,
SetbeforeSets), which could cause partial matches instead of capturing full keywords.
Here's the corrected code that fixes all these problems and outputs all matches as expected:
import re text_to_search = "Set Sets UnionShowSets jaja:={hi};" # Define patterns (order longer instruction keywords first to avoid partial matches) instruction_pattern = re.compile(r'SetUnion|SetIntersect|ShowSets|ShowSet|Sets|Set|Union|Intersect') operator_pattern = re.compile(r':=|{|}|;') id_pattern = re.compile(r'[a-zA-Z0-9]+') # Create a set of instruction keywords for quick filtering of IDs instruction_keywords = {"Set", "Sets", "ShowSet", "ShowSets", "Union", "Intersect", "SetUnion", "SetIntersect"} # Print all matching instructions print("Instructions:") for match in instruction_pattern.finditer(text_to_search): print(f"Instruction: {match}") # Print all matching operators print("\nOperators:") for match in operator_pattern.finditer(text_to_search): print(f"Operator: {match}") # Print all IDs (exclude instruction keywords) print("\nIDs:") for match in id_pattern.finditer(text_to_search): matched_text = match.group() if matched_text not in instruction_keywords: print(f"ID: {match}")
Output:
Instructions: Instruction: <re.Match object; span=(0, 3), match='Set'> Instruction: <re.Match object; span=(4, 8), match='Sets'> Instruction: <re.Match object; span=(8, 13), match='Union'> Instruction: <re.Match object; span=(13, 21), match='ShowSets'> Operators: Operator: <re.Match object; span=(26, 28), match=':='> Operator: <re.Match object; span=(28, 29), match='{'> Operator: <re.Match object; span=(31, 32), match='}'> Operator: <re.Match object; span=(32, 33), match=';'> IDs: ID: <re.Match object; span=(22, 26), match='jaja'> ID: <re.Match object; span=(29, 31), match='hi'>
Key Improvements:
re.finditer(): Iterates over all non-overlapping matches in the text, returning full match objects for each occurrence.- Ordered Instruction Patterns: Longer keywords (like
SetUnion,ShowSets) come first to ensure we capture full commands instead of partial matches (e.g.,Setinstead ofSets). - ID Filtering: We check each alphanumeric match against a set of instruction keywords to exclude commands from the ID results.
If you want to only match whole words (e.g., not extract Union from UnionShowSets), add word boundaries (\b) to the instruction pattern:
instruction_pattern = re.compile(r'\b(SetUnion|SetIntersect|ShowSets|ShowSet|Sets|Set|Union|Intersect)\b')
内容的提问来源于stack exchange,提问作者Sebas Silva
相关产品推荐
相关产品推荐

