如何用JavaScript提取.eml文件中script标签内JSON的指定值?
提取EML文件中JSON脚本的指定属性值
- Node.js环境无法直接使用浏览器的
document.getElementById等DOM API,需借助HTML解析库(如cheerio)处理HTML内容;另外EML是邮件格式,要先从中提取HTML正文,推荐用mailparser库完成解析。 - 先安装依赖:
npm install mailparser cheerio
具体实现步骤
读取并解析EML文件,提取HTML内容
用mailparser解析EML,拿到邮件的HTML正文:const fs = require('fs'); const { simpleParser } = require('mailparser'); const cheerio = require('cheerio'); // 读取本地EML文件 const emlContent = fs.readFileSync('./target-email.eml', 'utf8'); // 解析EML提取HTML正文 simpleParser(emlContent) .then(parsedMail => { const htmlContent = parsedMail.html; if (!htmlContent) { console.log('该邮件不含HTML内容'); return; } // 处理HTML中的目标script标签 extractTargetJson(htmlContent); }) .catch(err => console.error('解析EML失败:', err));定位目标script标签并解析JSON
用cheerio模拟DOM操作,筛选type="application/json"的script标签,提取内容后解析JSON并读取指定属性:function extractTargetJson(html) { const $ = cheerio.load(html); // 匹配type为application/json的script标签 const jsonScript = $('script[type="application/json"]'); if (jsonScript.length === 0) { console.log('未找到目标script标签'); return; } try { // 提取script内文本并解析为JSON对象 const jsonData = JSON.parse(jsonScript.text().trim()); // 获取指定属性,如publisher.name const publisherName = jsonData?.publisher?.name; publisherName ? console.log('提取到publisher.name:', publisherName) : console.log('未找到publisher.name属性'); } catch (parseErr) { console.error('JSON解析出错:', parseErr); } }
注意事项
- 确认EML的HTML部分确实存在
type="application/json"的script标签,且标签内内容是合法JSON(无语法错误)。 - 使用可选链运算符
?.可避免因属性不存在抛出异常,提升代码稳定性。
内容的提问来源于stack exchange,提问作者david
相关产品推荐
相关产品推荐

