JavaScript提取Reference标签多行内容:匹配至下一个星号
问题描述
用户有一段带星号标签的结构化文本,需要提取Reference:标签下的所有内容(直到遇到下一个*开头的标签为止),但现有JS代码仅能提取单行内容,无法满足需求。
原始文本
This is a code update * Official Name: Noner * Pub: https://content.upcodes.co/viewer/washington/wa-mechanical-code-2021 * Agency: Agency Ni * Reference: https://web.archive.org/web/20230226234118/https://lawfilesext.leg.wa.gov/law/wsr/agency/BuildingCodeCouncil.htm https://web.archive.org/web/20230303022030/https://lawfilesext.leg.wa.gov/law/wsr/2023/02/23-02-055.htm (#1) * Citation: WAC 51-52 / WSR 23-02-055 * Draft Doc Title: WSR 23-02-055 (#1) * Draft Source Doc: https://web.archive.org/web/20230303022030/https://lawfilesext.leg.wa.gov/law/wsr/2023/02/23-02-055.htm (#1) * Draft Drive: https://drive.google.com/file/d/1pYmwQS3t-ZX-Vyg9yBabtIpXZ7By2G6f/view?usp=share_link ( #1) * Final Doc Title: IECC Com Update(#1) IECC Res Update (#2) IECC Res Update (#3) * Final Source Doc: https://web.archive.org/web/20230303022130/https://apps.leg.wa.gov/wac/default.aspx?cite=51-52&full=true&pdf=true (#1) https://web.archive.org/web/20230303022030/https://lawfilesext.leg.wa.gov/law/wsr/2023/02/23-02-055.htm (#2) * Final Drive: https://web.archive.org/web/20230303022130/https://apps.leg.wa.gov/wac/default.aspx?cite=51-52&full=true&pdf=true (#1) https://web.archive.org/web/2023030302fdfdfg2130/https://apps.legfdg.gov/wac/default.aspx?cite=51-52&fdsfullfdsf=true&pfdsfdf=true (#2) * Effective Date: January 4, 2023
现有代码(仅支持单行提取)
//Extract Reference var reference = description.search("Reference:"); if(reference != -1){ reference = description.match(/(?<=^\* Reference\s*:)\s*[\n]*[^\n\r]*/m); reference = reference?.[0].trim(); }else{ reference = ''; } console.log('Reference: ' + reference);
期望输出
https://web.archive.org/web/20230226234118/https://lawfilesext.leg.wa.gov/law/wsr/agency/BuildingCodeCouncil.htm https://web.archive.org/web/20230303022030/https://lawfilesext.leg.wa.gov/law/wsr/2023/02/23-02-055.htm (#1)
解决方案
修改正则表达式,实现匹配Reference:标签后到下一个*开头标签前的所有内容,处理多余空白后即可得到目标结果:
//Extract Reference var reference = ''; const referenceMatch = description.match(/^\* Reference\s*:\s*([\s\S]*?)(?=\n\* |$)/m); if (referenceMatch) { reference = referenceMatch[1].trim(); } console.log('Reference: ' + reference);
代码说明
^\* Reference\s*:精准匹配开头的* Reference:标签,允许标签后带任意空白字符([\s\S]*?)非贪婪匹配任意字符(含换行),避免过度匹配到后续标签内容(?=\n\* |$)正向预查终止条件:匹配换行后接*的位置,或文本结尾,确保只提取到下一个标签前的内容- 最后用
trim()去除首尾多余的空白和换行符
内容的提问来源于stack exchange,提问作者alyssaeliyah
相关产品推荐
相关产品推荐

