仅匹配多行HTML<p>标签及内容的正则表达式需求
解决多行HTML
<p>标签的精准匹配问题 需求
需要从文档中仅匹配跨多行拆分的<p>标签,排除单行的<p>标签。现有正则<p>(.|\n)*?</p>会同时匹配单行和多行标签,无法满足需求。
测试用例
// 不应匹配(单行):
<p><b>start -</b> Uses an area level ("Name field" from levels.txt) to define where the player starts in the Act</p>// 应匹配(多行):
<p><b>maxnpcitemlevel -</b> Controls the maximum item level for items sold by the NPC in the Act </p>// 不应匹配(单行):
<p>This file controls global Act functionalities including item levels, monster behaviors, and waypoints</p>// 应匹配(多行):
<p><b>wanderingMonsterRegionTotal -</b> The maximum number of wandering monsters allowed at once </p>
解决方案
使用以下正则表达式,确保只匹配包含换行的多行<p>标签:
<p>(?:.|\n)*?\n(?:.|\n)*?</p>
正则逻辑说明
<p>:精准匹配起始标签(?:.|\n)*?:非贪婪匹配任意字符(含换行),避免过度匹配\n:强制匹配至少一个换行符,确保标签跨多行(?:.|\n)*?</p>:继续非贪婪匹配直到闭合标签</p>
若正则引擎支持DOTALL模式(让.匹配换行符),可使用更简洁的写法:
<p>.*?\n.*?</p>
内容的提问来源于stack exchange,提问作者WolfieeifloW
相关产品推荐
相关产品推荐

