如何使用JavaScript(Node.js)提取字符串中包含的第一个URL
Node.js 提取 HTML 字符串中首个 URL 的实现方法
下面提供两种常用实现方案,可根据你的场景选择:
方案1:正则匹配(无额外依赖,适合场景固定的情况)
如果你的输入字符串结构固定,URL 均为 http/https 协议且被引号包裹,可以直接用正则快速提取:
const inputStr = `<p> You left when I believed you would stay. You left my side when i needed you the most</p>**<img src="https://cloud-image.domain-name.com/storage/images/2021/picture_01.png" />** <br/> <p>Today, I don't have any bad words for you. I just want to thank you</p> **<img src="https://cloud-image.domain-name.com/storage/images/2021/picture_02.jpg" />`; // 匹配http/https开头,到最近的引号结束的内容 const urlReg = /https?:\/\/[^'"]+/i; const matchRes = inputStr.match(urlReg); const firstUrl = matchRes ? matchRes[0] : null; console.log(firstUrl); // 输出结果:https://cloud-image.domain-name.com/storage/images/2021/picture_01.png
方案2:HTML 解析库实现(更稳定,兼容复杂结构)
如果输入的 HTML 结构可能出现变化(比如 src 用单引号包裹、属性有多余空格等),可以用 Node.js 常用的 HTML 解析库 cheerio 避免正则匹配误差:
- 先安装依赖:
npm install cheerio - 实现代码:
const cheerio = require('cheerio'); const inputStr = `<p> You left when I believed you would stay. You left my side when i needed you the most</p>**<img src="https://cloud-image.domain-name.com/storage/images/2021/picture_01.png" />** <br/> <p>Today, I don't have any bad words for you. I just want to thank you</p> **<img src="https://cloud-image.domain-name.com/storage/images/2021/picture_02.jpg" />`; const $ = cheerio.load(inputStr); // 提取第一个img标签的src属性 const firstUrl = $('img').first().attr('src') || null; console.log(firstUrl); // 输出结果:https://cloud-image.domain-name.com/storage/images/2021/picture_01.png
内容的提问来源于stack exchange,提问作者Anna
相关产品推荐
相关产品推荐

