You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则表达式提取含内部HTML标签的锚标签内容?

Fixing Regex to Extract <a> Tag Content (Href, Inner HTML/Text)

Hey there! I see you're trying to pull the href value, inner text, and nested HTML from a <a class="related-article"> tag using regex, but your current pattern isn't hitting the mark. Let's walk through what's wrong with your existing regex, then fix it—and also talk about a more reliable alternative since regex isn't always the best fit for HTML.

What's Wrong With Your Current Regex?

Your pattern: (class="related-article"(?:\s|\n) )href="(. ?)"(>(.*?))" has a few key issues:

  • The (. ?) is a typo (missing the *), but even corrected to (.*?), it's not the most efficient way to capture the href value (it can accidentally match quotes if not careful).
  • The final (>(.*?))" part is incorrect—you're trying to match up to a quote, but the <a> tag closes with </a>, not a quote.
  • It doesn't account for flexible whitespace (like multiple spaces or line breaks) between attributes, which is common in real-world HTML.

Corrected Regex for Your Example

For the sample HTML you provided:

<a class="related-article" href="10.1182/blood-2017-11-812990"> <i>Blood</i> Commentary</a> on this article in this issue.</p>

Here's a regex that will reliably capture the href and inner content:

<a\s+class="related-article"\s+href="([^"]+)"\s*>([\s\S]*?)<\/a>

Breakdown of the Pattern:

  • <a\s+class="related-article"\s+href=": Matches the opening <a> tag, ensuring it has the related-article class, with flexible whitespace between attributes.
  • ([^"]+): Captures the href value—this matches all characters except a quote, which is safer than .*? because it stops exactly at the end of the href attribute.
  • "\s*>: Matches the closing quote of the href attribute, any trailing whitespace, and the tag's closing >.
  • ([\s\S]*?): Captures everything inside the <a> tag (including nested HTML like <i>). [\s\S] matches every character (including newlines), and *? makes it non-greedy so it stops at the first </a>.
  • <\/a>: Matches the closing </a> tag.

When you run this against your sample HTML, you'll get:

  • Capture Group 1: 10.1182/blood-2017-11-812990 (the href value)
  • Capture Group 2: <i>Blood</i> Commentary (the full inner content)

Handling Flexible Attribute Order

If the class and href attributes might appear in reverse order (e.g., <a href="..." class="related-article">), use this adjusted regex to cover both cases:

<a\s+(?:class="related-article"\s+href="([^"]+)"|href="([^"]+)"\s+class="related-article")\s*>([\s\S]*?)<\/a>

Here, you'll check which of the first two capture groups is non-empty to get the href value.

A Better Alternative: Use an HTML Parser

Regex works for simple cases, but HTML is a nested, flexible language—regex can break easily if the HTML structure changes (e.g., extra attributes, nested tags deeper than one level). For production code, use an HTML parsing library instead:

Example with Python's BeautifulSoup:

from bs4 import BeautifulSoup

html = '<a class="related-article" href="10.1182/blood-2017-11-812990"> <i>Blood</i> Commentary</a> on this article in this issue.</p>'
soup = BeautifulSoup(html, 'html.parser')

# Find the specific <a> tag
related_article = soup.find('a', class_='related-article')

if related_article:
    href = related_article['href']
    inner_html = related_article.decode_contents()  # Gets the raw inner HTML
    inner_text = related_article.get_text(strip=True)  # Gets just the text, stripped of tags

    print(f"Href: {href}")
    print(f"Inner HTML: {inner_html}")
    print(f"Inner Text: {inner_text}")

This will handle any valid HTML structure, no matter how messy it gets.

内容的提问来源于stack exchange,提问作者Surendran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:24:35