You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用正则表达式提取HTML标题与正文并忽略换行符的技术求助

正则提取与<body>内纯文本的优化方案 <a class="header-anchor" href="#正则提取与内纯文本的优化方案" aria-hidden="true">#</a></h1> <h2 id="提取标签内容">提取<title>标签内容 <a class="header-anchor" href="#提取标签内容" aria-hidden="true">#</a></h2> <p>标准<code><title></code>标签内不会嵌套其他HTML标签,用简单正则即可精准捕获:</p> <pre class="hljs"><code class="language-regex volc-pre-code">/<title>([^<]+)<\/title>/i </code></pre> <ul> <li><code>([^<]+)</code>:捕获标签间的标题文本,排除任何嵌套标签(符合规范的HTML不会出现此类情况)</li> <li><code>/i</code>:忽略大小写,兼容<code><TITLE></code>等非标准写法</li> </ul> <h2 id="提取并清理内的纯文本">提取并清理<body>内的纯文本 <a class="header-anchor" href="#提取并清理内的纯文本" aria-hidden="true">#</a></h2> <p>你原正则的问题是手动重复匹配逻辑,无法覆盖任意次数的标签/文本交替场景。推荐分两步处理,更简洁可靠:</p> <h3 id="_1-提取包裹的全部内容">1. 提取<body>包裹的全部内容 <a class="header-anchor" href="#_1-提取包裹的全部内容" aria-hidden="true">#</a></h3> <p>先捕获<code><body></code>到<code></body></code>之间的所有内容(包含标签):</p> <pre class="hljs"><code class="language-regex volc-pre-code">/<body>([\s\S]*?)<\/body>/i </code></pre> <ul> <li><code>[\s\S]*?</code>:匹配任意字符(含换行),非贪婪模式避免误匹配到后续的<code></body></code>(如果存在)</li> <li><code>/i</code>:忽略大小写</li> </ul> <h3 id="_2-移除所有html标签">2. 移除所有HTML标签 <a class="header-anchor" href="#_2-移除所有html标签" aria-hidden="true">#</a></h3> <p>对提取到的body内容,用全局替换清除所有标签:</p> <pre class="hljs"><code class="language-regex volc-pre-code">/<[^>]+>/g </code></pre> <ul> <li><code><[^>]+></code>:匹配任意HTML标签(含带属性的标签,如<code><a href="xxx"></code>)</li> <li><code>/g</code>:全局替换,移除所有匹配项</li> </ul> <h3 id="代码示例(python)">代码示例(Python) <a class="header-anchor" href="#代码示例(python)" aria-hidden="true">#</a></h3> <pre class="hljs"><code class="language-python volc-pre-code">import re sample_html = """<html> <head><title>Some title</title></head> <body>Here<p> is some </p>content <a href="www.somesite.com"> click</body> </html>""" # 提取标题 title = re.search(r'<title>([^<]+)</title>', sample_html, re.IGNORECASE).group(1) # 提取并清理正文 body_raw = re.search(r'<body>([\s\S]*?)</body>', sample_html, re.IGNORECASE).group(1) clean_body = re.sub(r'<[^>]+>', '', body_raw).strip() print(f"Title: {title}") print(f"Clean Body: {clean_body}") </code></pre> <p>输出:</p> <pre class="hljs"><code class="language- volc-pre-code">Title: Some title Clean Body: Here is some content click </code></pre> <h2 id="原正则的优化方向">原正则的优化方向 <a class="header-anchor" href="#原正则的优化方向" aria-hidden="true">#</a></h2> <p>你原正则的核心逻辑是<code>([^<]*)(?:<[^>]*+>)*</code>,手动重复四次无法覆盖所有场景,可改用重复分组替代:</p> <pre class="hljs"><code class="language-regex volc-pre-code">/<body>(?:([^<]*)<[^>]*>)*([^<]*)<\/body>/i </code></pre> <p>但这种方式的捕获组会被多次覆盖,需要遍历所有捕获片段拼接,效率不如先提取整体再清理标签。</p> <hr> <p>内容的提问来源于stack exchange,提问作者Ersin Nurtin</p>
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 16:16:06