使用正则表达式提取HTML标题与正文并忽略换行符的技术求助
正则提取与<body>内纯文本的优化方案 <a class="header-anchor" href="#正则提取与内纯文本的优化方案" aria-hidden="true">#</a></h1>
<h2 id="提取标签内容">提取<title>标签内容 <a class="header-anchor" href="#提取标签内容" aria-hidden="true">#</a></h2>
<p>标准<code><title></code>标签内不会嵌套其他HTML标签,用简单正则即可精准捕获:</p>
<pre class="hljs"><code class="language-regex volc-pre-code">/<title>([^<]+)<\/title>/i
</code></pre>
<ul>
<li><code>([^<]+)</code>:捕获标签间的标题文本,排除任何嵌套标签(符合规范的HTML不会出现此类情况)</li>
<li><code>/i</code>:忽略大小写,兼容<code><TITLE></code>等非标准写法</li>
</ul>
<h2 id="提取并清理内的纯文本">提取并清理<body>内的纯文本 <a class="header-anchor" href="#提取并清理内的纯文本" aria-hidden="true">#</a></h2>
<p>你原正则的问题是手动重复匹配逻辑,无法覆盖任意次数的标签/文本交替场景。推荐分两步处理,更简洁可靠:</p>
<h3 id="_1-提取包裹的全部内容">1. 提取<body>包裹的全部内容 <a class="header-anchor" href="#_1-提取包裹的全部内容" aria-hidden="true">#</a></h3>
<p>先捕获<code><body></code>到<code></body></code>之间的所有内容(包含标签):</p>
<pre class="hljs"><code class="language-regex volc-pre-code">/<body>([\s\S]*?)<\/body>/i
</code></pre>
<ul>
<li><code>[\s\S]*?</code>:匹配任意字符(含换行),非贪婪模式避免误匹配到后续的<code></body></code>(如果存在)</li>
<li><code>/i</code>:忽略大小写</li>
</ul>
<h3 id="_2-移除所有html标签">2. 移除所有HTML标签 <a class="header-anchor" href="#_2-移除所有html标签" aria-hidden="true">#</a></h3>
<p>对提取到的body内容,用全局替换清除所有标签:</p>
<pre class="hljs"><code class="language-regex volc-pre-code">/<[^>]+>/g
</code></pre>
<ul>
<li><code><[^>]+></code>:匹配任意HTML标签(含带属性的标签,如<code><a href="xxx"></code>)</li>
<li><code>/g</code>:全局替换,移除所有匹配项</li>
</ul>
<h3 id="代码示例(python)">代码示例(Python) <a class="header-anchor" href="#代码示例(python)" aria-hidden="true">#</a></h3>
<pre class="hljs"><code class="language-python volc-pre-code">import re
sample_html = """<html>
<head><title>Some title</title></head>
<body>Here<p> is some </p>content <a href="www.somesite.com">
click</body>
</html>"""
# 提取标题
title = re.search(r'<title>([^<]+)</title>', sample_html, re.IGNORECASE).group(1)
# 提取并清理正文
body_raw = re.search(r'<body>([\s\S]*?)</body>', sample_html, re.IGNORECASE).group(1)
clean_body = re.sub(r'<[^>]+>', '', body_raw).strip()
print(f"Title: {title}")
print(f"Clean Body: {clean_body}")
</code></pre>
<p>输出:</p>
<pre class="hljs"><code class="language- volc-pre-code">Title: Some title
Clean Body: Here is some content click
</code></pre>
<h2 id="原正则的优化方向">原正则的优化方向 <a class="header-anchor" href="#原正则的优化方向" aria-hidden="true">#</a></h2>
<p>你原正则的核心逻辑是<code>([^<]*)(?:<[^>]*+>)*</code>,手动重复四次无法覆盖所有场景,可改用重复分组替代:</p>
<pre class="hljs"><code class="language-regex volc-pre-code">/<body>(?:([^<]*)<[^>]*>)*([^<]*)<\/body>/i
</code></pre>
<p>但这种方式的捕获组会被多次覆盖,需要遍历所有捕获片段拼接,效率不如先提取整体再清理标签。</p>
<hr>
<p>内容的提问来源于stack exchange,提问作者Ersin Nurtin</p>
相关产品推荐
相关产品推荐

