You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现类Flex的正则匹配,修正指令/ID/运算符识别问题

Solution

Your original code had three main issues:

  1. re.search() only returns the first match, so you missed all subsequent matches for instructions, operators, and IDs.
  2. The ID regex didn't exclude instruction keywords, leading to false matches on commands like Set.
  3. The instruction regex ordered shorter patterns first (e.g., Set before Sets), which could cause partial matches instead of capturing full keywords.

Here's the corrected code that fixes all these problems and outputs all matches as expected:

import re

text_to_search = "Set Sets UnionShowSets jaja:={hi};"

# Define patterns (order longer instruction keywords first to avoid partial matches)
instruction_pattern = re.compile(r'SetUnion|SetIntersect|ShowSets|ShowSet|Sets|Set|Union|Intersect')
operator_pattern = re.compile(r':=|{|}|;')
id_pattern = re.compile(r'[a-zA-Z0-9]+')

# Create a set of instruction keywords for quick filtering of IDs
instruction_keywords = {"Set", "Sets", "ShowSet", "ShowSets", "Union", "Intersect", "SetUnion", "SetIntersect"}

# Print all matching instructions
print("Instructions:")
for match in instruction_pattern.finditer(text_to_search):
    print(f"Instruction: {match}")

# Print all matching operators
print("\nOperators:")
for match in operator_pattern.finditer(text_to_search):
    print(f"Operator: {match}")

# Print all IDs (exclude instruction keywords)
print("\nIDs:")
for match in id_pattern.finditer(text_to_search):
    matched_text = match.group()
    if matched_text not in instruction_keywords:
        print(f"ID: {match}")

Output:

Instructions:
Instruction: <re.Match object; span=(0, 3), match='Set'>
Instruction: <re.Match object; span=(4, 8), match='Sets'>
Instruction: <re.Match object; span=(8, 13), match='Union'>
Instruction: <re.Match object; span=(13, 21), match='ShowSets'>

Operators:
Operator: <re.Match object; span=(26, 28), match=':='>
Operator: <re.Match object; span=(28, 29), match='{'>
Operator: <re.Match object; span=(31, 32), match='}'>
Operator: <re.Match object; span=(32, 33), match=';'>

IDs:
ID: <re.Match object; span=(22, 26), match='jaja'>
ID: <re.Match object; span=(29, 31), match='hi'>

Key Improvements:

  • re.finditer(): Iterates over all non-overlapping matches in the text, returning full match objects for each occurrence.
  • Ordered Instruction Patterns: Longer keywords (like SetUnion, ShowSets) come first to ensure we capture full commands instead of partial matches (e.g., Set instead of Sets).
  • ID Filtering: We check each alphanumeric match against a set of instruction keywords to exclude commands from the ID results.

If you want to only match whole words (e.g., not extract Union from UnionShowSets), add word boundaries (\b) to the instruction pattern:

instruction_pattern = re.compile(r'\b(SetUnion|SetIntersect|ShowSets|ShowSet|Sets|Set|Union|Intersect)\b')

内容的提问来源于stack exchange,提问作者Sebas Silva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:55:35