如何避免UrlFetchApp.fetch无限阻塞脚本执行?
问题分析与解决思路
核心原因
这种阻塞到脚本最大执行时间才报错的情况,本质是目标服务器故意挂起连接不返回响应,而非普通的网络超时:
- 反爬机制触发:像usnews这类网站会检测请求的来源特征,UrlFetchApp默认的请求头(无浏览器标识、缺少常见字段)会被识别为爬虫,服务器不会直接拒绝请求,而是保持连接不发送任何数据,让脚本一直等待直到耗尽执行时间。
- 服务器资源限制:部分服务器在负载过高时,会对新请求保持长连接但不处理,直到连接被强制断开。
处理方案
1. 强制设置请求超时
UrlFetchApp默认单请求超时为60秒,但如果服务器一直不返回,脚本会持续等待直到达到整个脚本的最大执行时间(6分钟)。必须手动设置timeoutMillis参数,让请求在指定时间内无响应就主动终止:
function fetch_test() { try { const options = { timeoutMillis: 10000 // 10秒超时 }; const res = UrlFetchApp.fetch('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands', options); Logger.log(res.getContentText()); } catch (e) { Logger.log("请求超时或失败:" + e.message); } }
2. 模拟浏览器请求头
添加常见的浏览器请求头,规避反爬机制的识别:
function fetch_test() { try { const options = { timeoutMillis: 10000, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' } }; const res = UrlFetchApp.fetch('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands', options); Logger.log(res.getContentText()); } catch (e) { Logger.log("请求超时或失败:" + e.message); } }
3. 实现有限重试机制
针对偶尔出现的阻塞情况,可在超时后重试有限次数,避免单次失败就终止流程:
function fetchWithRetry(url, maxRetries = 2, timeout = 10000) { const options = { timeoutMillis: timeout, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } }; for (let i = 0; i <= maxRetries; i++) { try { return UrlFetchApp.fetch(url, options); } catch (e) { if (i === maxRetries) { throw new Error(`重试${maxRetries}次后仍失败:${e.message}`); } // 重试前短暂延迟,降低被反爬识别的概率 Utilities.sleep(1000); } } } function fetch_test() { try { const res = fetchWithRetry('https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands'); Logger.log(res.getContentText()); } catch (e) { Logger.log(e.message); } }
4. fetchAll批量处理的注意事项
使用fetchAll时,同样要给每个请求配置timeoutMillis,避免单个阻塞请求拖垮整个批量任务:
function fetchAll_test() { const requests = [ { url: 'https://www.usnews.com/news/world/articles/2022-12-06/turkey-again-threatens-greece-for-arming-aegean-islands', timeoutMillis: 10000, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } }, // 其他请求... ]; try { const responses = UrlFetchApp.fetchAll(requests); responses.forEach(res => Logger.log(res.getResponseCode())); } catch (e) { Logger.log("批量请求失败:" + e.message); } }
关键注意点
- 永远不要依赖UrlFetchApp的默认超时,必须手动设置合理的
timeoutMillis(建议5-15秒),避免脚本被无意义的长连接耗尽执行时间。 - 反爬机制会持续更新,若模拟浏览器头仍无效,可尝试添加
Accept-Language、Referer等更多常见请求头,但不要伪造违规标识。 - 重试次数不宜过多,避免给目标服务器造成压力,同时添加短暂延迟,降低被识别为恶意爬虫的概率。
内容的提问来源于stack exchange,提问作者Costas Kirgoussios
相关产品推荐
相关产品推荐

