如何用正则表达式提取含内部HTML标签的锚标签内容?
<a> Tag Content (Href, Inner HTML/Text) Hey there! I see you're trying to pull the href value, inner text, and nested HTML from a <a class="related-article"> tag using regex, but your current pattern isn't hitting the mark. Let's walk through what's wrong with your existing regex, then fix it—and also talk about a more reliable alternative since regex isn't always the best fit for HTML.
What's Wrong With Your Current Regex?
Your pattern: (class="related-article"(?:\s|\n) )href="(. ?)"(>(.*?))" has a few key issues:
- The
(. ?)is a typo (missing the*), but even corrected to(.*?), it's not the most efficient way to capture the href value (it can accidentally match quotes if not careful). - The final
(>(.*?))"part is incorrect—you're trying to match up to a quote, but the<a>tag closes with</a>, not a quote. - It doesn't account for flexible whitespace (like multiple spaces or line breaks) between attributes, which is common in real-world HTML.
Corrected Regex for Your Example
For the sample HTML you provided:
<a class="related-article" href="10.1182/blood-2017-11-812990"> <i>Blood</i> Commentary</a> on this article in this issue.</p>
Here's a regex that will reliably capture the href and inner content:
<a\s+class="related-article"\s+href="([^"]+)"\s*>([\s\S]*?)<\/a>
Breakdown of the Pattern:
<a\s+class="related-article"\s+href=": Matches the opening<a>tag, ensuring it has therelated-articleclass, with flexible whitespace between attributes.([^"]+): Captures the href value—this matches all characters except a quote, which is safer than.*?because it stops exactly at the end of the href attribute."\s*>: Matches the closing quote of the href attribute, any trailing whitespace, and the tag's closing>.([\s\S]*?): Captures everything inside the<a>tag (including nested HTML like<i>).[\s\S]matches every character (including newlines), and*?makes it non-greedy so it stops at the first</a>.<\/a>: Matches the closing</a>tag.
When you run this against your sample HTML, you'll get:
- Capture Group 1:
10.1182/blood-2017-11-812990(the href value) - Capture Group 2:
<i>Blood</i> Commentary(the full inner content)
Handling Flexible Attribute Order
If the class and href attributes might appear in reverse order (e.g., <a href="..." class="related-article">), use this adjusted regex to cover both cases:
<a\s+(?:class="related-article"\s+href="([^"]+)"|href="([^"]+)"\s+class="related-article")\s*>([\s\S]*?)<\/a>
Here, you'll check which of the first two capture groups is non-empty to get the href value.
A Better Alternative: Use an HTML Parser
Regex works for simple cases, but HTML is a nested, flexible language—regex can break easily if the HTML structure changes (e.g., extra attributes, nested tags deeper than one level). For production code, use an HTML parsing library instead:
Example with Python's BeautifulSoup:
from bs4 import BeautifulSoup html = '<a class="related-article" href="10.1182/blood-2017-11-812990"> <i>Blood</i> Commentary</a> on this article in this issue.</p>' soup = BeautifulSoup(html, 'html.parser') # Find the specific <a> tag related_article = soup.find('a', class_='related-article') if related_article: href = related_article['href'] inner_html = related_article.decode_contents() # Gets the raw inner HTML inner_text = related_article.get_text(strip=True) # Gets just the text, stripped of tags print(f"Href: {href}") print(f"Inner HTML: {inner_html}") print(f"Inner Text: {inner_text}")
This will handle any valid HTML structure, no matter how messy it gets.
内容的提问来源于stack exchange,提问作者Surendran

