如何用正则表达式提取HTML中script标签的src值与内容?
正则提取HTML中script标签的src属性与内部代码
需求
给定包含HTML的字符串,需要用两个正则分别完成:
- 提取script标签的src属性值
- 提取script标签内部的代码内容
原始HTML字符串示例:
let htmlTEXT = `<h1>page 1</h1> <input type="text" spellcheck="false" data-ms-editor="true"> <script src="/static/test.js"> console.log("hello"); console.log("goodbye"); </script>`
期望结果:
- src属性值:
/static/test.js - 内部代码:
console.log("hello"); console.log("goodbye");
解决方案
1. 提取src属性值
使用正则匹配带src属性的script标签,捕获属性值:
const srcRegex = /<script\s+src="([^"]+)"/i; const srcMatch = htmlTEXT.match(srcRegex); const srcValue = srcMatch ? srcMatch[1] : ''; console.log(srcValue); // 输出: /static/test.js
正则说明:
<script\s+src=":匹配script标签开头,后跟至少一个空格和src="([^"]+):捕获双引号内的所有非双引号字符,即src属性值i:忽略大小写,兼容<SCRIPT>这类大写标签
2. 提取script内部代码
使用正则捕获script开闭标签之间的内容:
const contentRegex = /<script[^>]*>([\s\S]*?)<\/script>/i; const contentMatch = htmlTEXT.match(contentRegex); const scriptContent = contentMatch ? contentMatch[1].trim() : ''; console.log(scriptContent);
正则说明:
<script[^>]*>:匹配script标签开头,[^>]*匹配标签内所有属性内容(直到>`结束)([\s\S]*?):捕获标签内所有内容,[\s\S]匹配包括换行在内的任意字符,*?是非贪婪匹配,避免匹配多个script标签时的溢出<\/script>:匹配闭合的script标签trim():去除内容前后的空白换行,让结果更整洁
注意事项
- 上述正则针对示例中的双引号属性、单个script标签场景设计,若遇到src用单引号/无引号、多个script标签的情况,需调整正则逻辑(比如用
['"]?([^'"]+)['"]?兼容引号类型,用matchAll处理多标签) - 正则处理HTML存在局限性,复杂场景建议使用DOM解析方案(比如浏览器端的
DOMParser)
内容的提问来源于stack exchange,提问作者Sean Lee
相关产品推荐
相关产品推荐

