You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Google Apps Script抓取网页时正确解析特殊字符

Google Apps Script 抓取网页HTML实体解码解决方案

问题根因

你遇到的’这类格式的字符属于HTML实体编码,该转义操作是目标网站返回响应内容时自带的,getContentText()方法仅原样返回HTTP响应的文本内容,没有额外执行编码逻辑。

通用解决方案

方案1:内置XmlService解码(推荐)

利用Google Apps Script内置的XML解析能力自动解码所有标准HTML实体,无需手动维护映射规则,适配绝大多数场景:

// 封装通用解码方法
function decodeHtmlEntities(html) {
  if (!html) return html;
  const parseDoc = XmlService.parse(`<root>${html}</root>`);
  return parseDoc.getRootElement().getText();
}

// 原有逻辑改造
var response = UrlFetchApp.fetch("https://podtail.com/de/top-podcasts/de/");
var rawHtml = response.getContentText();
var decodedHtml = decodeHtmlEntities(rawHtml); // 此处已完成所有特殊字符解码

方案2:自定义映射兼容解码

如果目标页面存在大量不符合XML规范的内容,导致XmlService解析报错,可以使用自定义实体映射的方式适配:

function decodeHtmlEntitiesFallback(html) {
  // 可按需补充需要匹配的实体规则
  const entityMap = {
    '&amp;': '&',
    '&lt;': '<',
    '&gt;': '>',
    '&quot;': '"',
    '&apos;': "'",
    '&rsquo;': '’',
    '&lsquo;': '‘',
    '&ldquo;': '“',
    '&rdquo;': '”',
    '&nbsp;': ' '
  };
  return html.replace(/&[a-zA-Z0-9#]+;/g, match => entityMap[match] || match);
}

内容的提问来源于stack exchange,提问作者noah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 19:24:01