正则多结果匹配需求:提取title类名及子li的name与src属性值
Hey there! Let's break down how to solve your problem with regular expressions. First, I'll assume you're working with HTML-like content where elements with the title class are followed by a <ul> containing your target <li> elements (since you mentioned child li elements). Let's start with a sample input to ground our solution:
<div class="title">Fruits Collection</div> <ul> <li name="red_apple" src="/assets/apple-red.png">Crisp Red Apple</li> <li name="yellow_banana" src="/assets/banana-yellow.png">Sweet Yellow Banana</li> </ul> <div class="title">Vegetable Picks</div> <ul> <li name="orange_carrot" src="/assets/carrot-orange.png">Crunchy Orange Carrot</li> <li name="green_spinach" src="/assets/spinach-green.png">Nutritious Green Spinach</li> <li name="purple_eggplant" src="/assets/eggplant-purple.png">Rich Purple Eggplant</li> </ul>
Step 1: The Regex Pattern
We'll use a pattern that first captures the text inside the title class element, then extracts the name and src attributes from every child <li>. Here's the pattern (tweaked for flexibility with whitespace):
<div class="title">(.*?)<\/div>\s*<ul>\s*(?:<li name="(.*?)" src="(.*?)".*?<\/li>\s*)*<\/ul>
Let's break down what each part does:
<div class="title">(.*?)<\/div>: Captures the text inside the title element (group 1) — the.*?is a non-greedy match to stop at the closing</div>.\s*: Matches any amount of whitespace (newlines, spaces, tabs) to handle messy formatting.<ul>\s*: Matches the opening unordered list tag and any following whitespace.(?:<li name="(.*?)" src="(.*?)".*?<\/li>\s*)*: A non-capturing group that repeats for every<li>:name="(.*?)": Captures thenameattribute value (group 2).src="(.*?)": Captures thesrcattribute value (group 3)..*?<\/li>: Matches the rest of the li content until the closing tag.
Step 2: Implementing the Pattern (Python Example)
Since regex engines handle repeated capture groups differently, here's a practical Python implementation that extracts all title-name-src pairs correctly:
import re # Your input content content = """ <div class="title">Fruits Collection</div> <ul> <li name="red_apple" src="/assets/apple-red.png">Crisp Red Apple</li> <li name="yellow_banana" src="/assets/banana-yellow.png">Sweet Yellow Banana</li> </ul> <div class="title">Vegetable Picks</div> <ul> <li name="orange_carrot" src="/assets/carrot-orange.png">Crunchy Orange Carrot</li> <li name="green_spinach" src="/assets/spinach-green.png">Nutritious Green Spinach</li> <li name="purple_eggplant" src="/assets/eggplant-purple.png">Rich Purple Eggplant</li> </ul> """ # First, match each title + ul block title_block_pattern = re.compile(r'<div class="title">(.*?)<\/div>\s*<ul>(.*?)<\/ul>', re.DOTALL) for match in title_block_pattern.finditer(content): title_name = match.group(1).strip() print(f"📋 Title: {title_name}") # Now extract li name and src from the captured ul content li_pattern = re.compile(r'<li name="(.*?)" src="(.*?)".*?<\/li>') li_matches = li_pattern.findall(match.group(2)) for name, src in li_matches: print(f" - Name: *{name}*, Src: `{src}`")
Step 3: Key Notes & Caveats
- Format Flexibility: If your HTML uses single quotes for attributes (e.g.,
name='apple'), adjust the regex to use'instead of". - Tag Variations: If your
titleclass is on a different tag (like<h2>instead of<div>), update the pattern's opening/closing tags. - Regex Limitation: Regular expressions aren't designed for nested HTML structures. If your
<li>elements contain nested tags or the HTML is heavily malformed, you'll get better results with an HTML parser like Python'sBeautifulSoupor JavaScript'sCheerio.
内容的提问来源于stack exchange,提问作者MGR

