正则表达式需求:匹配开始标签,排除闭合标签后的开始标签
Got it, I understand exactly what you're going for here—you want to match every opening tag like <any-string>, but exclude any instance of that same tag that shows up after its corresponding closing tag </any-string>. It's just like matching foo but skipping any foo that comes right after bar.
Let's break down how to solve this, since regex alone has some limits here, but we've got solid options depending on your use case:
Option 1: Use code for full control (most reliable)
Regex is great for pattern matching, but it's stateless—it can't keep track of "which tags have been closed already" as it scans the string. For the most accurate results, especially if you have complex content with multiple tags or gaps between closing and opening tags, use a simple script to track closed tags:
Example in Python:
import re def get_valid_opening_tags(input_str): closed_tags = set() valid_tags = [] # Match all opening and closing tags in the string for tag_match in re.finditer(r'</?(\w+)>', input_str): tag_name = tag_match.group(1) full_tag = tag_match.group(0) if full_tag.startswith('</'): # Add this tag to our "already closed" set closed_tags.add(tag_name) else: # Only keep the opening tag if it hasn't been closed yet if tag_name not in closed_tags: valid_tags.append(full_tag) return valid_tags # Test your sample scenarios: print(get_valid_opening_tags("<h1>")) # Output: ['<h1>'] print(get_valid_opening_tags("<h1> abc </h1>")) # Output: ['<h1>'] print(get_valid_opening_tags("<abc>")) # Output: ['<abc>'] print(get_valid_opening_tags("</h1><h1>")) # Output: [] (excludes the second <h1>)
This works by maintaining a set of tags we've seen closed. Every time we hit an opening tag, we check if it's already been closed—if not, we count it as a valid match.
Option 2: Regex (for simple scenarios only)
If your use case is super straightforward (e.g., you only need to exclude opening tags that directly follow their closing tag with no content in between), you can use a regex with PCRE's (*SKIP)(*FAIL) mechanism to skip those invalid tags:
</(\w+)>\s*<\1>(*SKIP)(*FAIL)|<(\w+)>
What this does:
- The first part
</(\w+)>\s*<\1>(*SKIP)(*FAIL)matches any closing tag followed immediately (with optional whitespace) by its corresponding opening tag, then tells the regex engine to skip and fail these matches. - The second part
<(\w+)>matches all other opening tags.
Important caveat: This won't work if there's any content between the closing tag and the invalid opening tag (like </h1> some text <h1>)—it will still match that second <h1>. That's why the code approach is better for most real-world cases.
Testing your scenarios
Both methods will correctly handle your examples:
- Matches
<h1>in<h1> - Matches
<h1>in<h1> abc </h1> - Matches
<abc>in<abc> - Excludes
<h1>in</h1><h1>
内容的提问来源于stack exchange,提问作者sql_dummy

