Node.js+Axios构建爬虫遇connect ETIMEDOUT错误求助
解决Axios爬取Emojipedia详情页时的ETIMEDOUT问题
你遇到的问题核心是批量请求触发了网站限制,或者Keep-Alive配置没真正生效,导致连接超时。下面是针对性的修复方案:
可能的触发原因
- 一次性发起大量详情页请求,并发过高,被网站反爬机制拦截,服务器拒绝建立新连接
- 每次请求都新建HTTP Agent,Keep-Alive没复用连接,浪费资源且容易被限流
- 仅设置了Axios总超时,未单独配置连接阶段的超时,导致连接等待过久超时
具体修复步骤
1. 控制请求频率,限制并发数
别一次性爬所有详情页,要么串行请求加随机间隔,要么用库限制并发:
const http = require('http'); const https = require('https'); const axios = require('axios'); const cheerio = require('cheerio'); // 随机延时函数,避免固定间隔被识别 const delay = (ms) => new Promise(resolve => setTimeout(resolve, ms)); async function crawlEmojis() { // 先爬首页拿表情列表 const homeRes = await axios.get('https://emojipedia.org/', { httpAgent: new http.Agent({ keepAlive: true }), httpsAgent: new https.Agent({ keepAlive: true }), timeout: 60000, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/118.0.0.0 Safari/537.36' } }); const $ = cheerio.load(homeRes.data); const emojiList = $('.emoji-list li').map((i, el) => ({ name: $(el).find('a').attr('title'), url: 'https://emojipedia.org' + $(el).find('a').attr('href') })).get(); // 串行遍历,每个请求间隔1-3秒随机延时 for (const emoji of emojiList) { try { await delay(Math.random() * 2000 + 1000); const detailRes = await axios.get(emoji.url, { httpAgent: new http.Agent({ keepAlive: true }), httpsAgent: new https.Agent({ keepAlive: true }), timeout: 60000, headers: { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) Chrome/118.0.0.0 Safari/537.36' } }); const detail$ = cheerio.load(detailRes.data); emoji.description = detail$('.description').text().trim(); console.log(`搞定:${emoji.name}`); } catch (err) { console.error(`爬${emoji.url}挂了:`, err.message); } } } crawlEmojis();
2. 复用Agent实例,让Keep-Alive真正生效
别每次请求都新建Agent,全局创建一次复用,这样能保持长连接:
// 全局创建可复用的Agent const httpAgent = new http.Agent({ keepAlive: true, keepAliveMsecs: 30000, // 连接保持30秒 maxSockets: 5 // 限制最大并发连接数,别搞太猛 }); const httpsAgent = new https.Agent({ keepAlive: true, keepAliveMsecs: 30000, maxSockets: 5 }); // 所有请求都用这两个Agent const homeRes = await axios.get('https://emojipedia.org/', { httpAgent, httpsAgent, timeout: 60000, headers: { 'User-Agent': '你的浏览器UA' } }); // 详情页请求同样复用 const detailRes = await axios.get(emoji.url, { httpAgent, httpsAgent, timeout: 60000, headers: { 'User-Agent': '你的浏览器UA' } });
3. 细化超时配置,拆分连接和响应超时
Axios的timeout是总超时,给Agent单独设置连接超时,避免连接阶段卡太久:
const httpsAgent = new https.Agent({ keepAlive: true, timeout: 15000, // 连接超时15秒,比总超时短 keepAliveMsecs: 30000, maxSockets: 5 }); // 总请求超时还是60秒(包括响应时间) axios.get(emoji.url, { httpsAgent, timeout: 60000, headers: { 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5' } });
4. 加个重试机制,对付偶发超时
用axios-retry库自动重试超时请求:
const axiosRetry = require('axios-retry'); // 配置重试规则:只重试超时错误,最多3次,间隔递增 axiosRetry(axios, { retries: 3, retryDelay: (count) => count * 2000, retryCondition: (err) => err.code === 'ETIMEDOUT' });
额外提醒
- 一定要用真实的浏览器User-Agent,别用默认的axios标识
- 遵守网站的robots.txt,Emojipedia允许爬取但别太激进
- 把失败的URL记录下来,后续单独补爬
内容的提问来源于stack exchange,提问作者Anish-Aby
相关产品推荐
相关产品推荐

