You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免UrlFetchApp.fetch无限阻塞脚本执行?

问题分析与解决思路

核心原因

这种阻塞到脚本最大执行时间才报错的情况,本质是目标服务器故意挂起连接不返回响应,而非普通的网络超时:

  • 反爬机制触发:像usnews这类网站会检测请求的来源特征,UrlFetchApp默认的请求头(无浏览器标识、缺少常见字段)会被识别为爬虫,服务器不会直接拒绝请求,而是保持连接不发送任何数据,让脚本一直等待直到耗尽执行时间。
  • 服务器资源限制:部分服务器在负载过高时,会对新请求保持长连接但不处理,直到连接被强制断开。

处理方案

1. 强制设置请求超时

UrlFetchApp默认单请求超时为60秒,但如果服务器一直不返回,脚本会持续等待直到达到整个脚本的最大执行时间(6分钟)。必须手动设置timeoutMillis参数,让请求在指定时间内无响应就主动终止:

function fetch_test() {
  try {
    const options = {
      timeoutMillis: 10000 // 10秒超时
    };
    const res = UrlFetchApp.fetch('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands', options);
    Logger.log(res.getContentText());
  } catch (e) {
    Logger.log("请求超时或失败:" + e.message);
  }
}

2. 模拟浏览器请求头

添加常见的浏览器请求头,规避反爬机制的识别:

function fetch_test() {
  try {
    const options = {
      timeoutMillis: 10000,
      headers: {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'
      }
    };
    const res = UrlFetchApp.fetch('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands', options);
    Logger.log(res.getContentText());
  } catch (e) {
    Logger.log("请求超时或失败:" + e.message);
  }
}

3. 实现有限重试机制

针对偶尔出现的阻塞情况,可在超时后重试有限次数,避免单次失败就终止流程:

function fetchWithRetry(url, maxRetries = 2, timeout = 10000) {
  const options = {
    timeoutMillis: timeout,
    headers: {
      'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
  };
  
  for (let i = 0; i <= maxRetries; i++) {
    try {
      return UrlFetchApp.fetch(url, options);
    } catch (e) {
      if (i === maxRetries) {
        throw new Error(`重试${maxRetries}次后仍失败:${e.message}`);
      }
      // 重试前短暂延迟,降低被反爬识别的概率
      Utilities.sleep(1000);
    }
  }
}

function fetch_test() {
  try {
    const res = fetchWithRetry('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands');
    Logger.log(res.getContentText());
  } catch (e) {
    Logger.log(e.message);
  }
}

4. fetchAll批量处理的注意事项

使用fetchAll时,同样要给每个请求配置timeoutMillis,避免单个阻塞请求拖垮整个批量任务:

function fetchAll_test() {
  const requests = [
    {
      url: 'https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands',
      timeoutMillis: 10000,
      headers: {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
      }
    },
    // 其他请求...
  ];
  
  try {
    const responses = UrlFetchApp.fetchAll(requests);
    responses.forEach(res => Logger.log(res.getResponseCode()));
  } catch (e) {
    Logger.log("批量请求失败:" + e.message);
  }
}

关键注意点

  • 永远不要依赖UrlFetchApp的默认超时,必须手动设置合理的timeoutMillis(建议5-15秒),避免脚本被无意义的长连接耗尽执行时间。
  • 反爬机制会持续更新,若模拟浏览器头仍无效,可尝试添加Accept-Language、Referer等更多常见请求头,但不要伪造违规标识。
  • 重试次数不宜过多,避免给目标服务器造成压力,同时添加短暂延迟,降低被识别为恶意爬虫的概率。

内容的提问来源于stack exchange,提问作者Costas Kirgoussios

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 03:31:05