如何在axios+cheerio实现的getLinks()函数中添加Yellowpages下一页抓取逻辑
要实现自动翻页持续抓取,只需要在getLinks函数处理完当前页店铺数据后,增加下一页链接解析和递归调用的逻辑即可,修改方案如下:
核心逻辑说明
- Yellowpages搜索结果页的下一页按钮固定类名为
next,定位该元素即可拿到下一页的相对路径 - 拼接host得到完整下一页地址后,递归调用
getLinks即可自动处理后续所有页面 - 没有匹配到下一页元素时就停止抓取,自动结束流程
修改后完整代码
const axios = require('axios'); const cheerio = require('cheerio'); const startUrl = 'https://www.yellowpages.com/search?search_terms=Pizza&geo_location_terms=San+Francisco%2C+CA'; const host = 'https://www.yellowpages.com'; const getLinks = async (url,host,callback) => { const { data } = await axios.get(url); const $ = cheerio.load(data); // 原有提取当前页店铺链接逻辑不变 $('[class="result"] a.business-name').each(function(){ let items = $(this).find('span').text(); let links = host + $(this).attr("href"); callback(items,links); }); // 新增下一页解析和调用逻辑 const nextPageHref = $('.pagination a.next').attr('href'); if (nextPageHref) { const nextPageUrl = host + nextPageHref; // 递归调用处理下一页 await getLinks(nextPageUrl, host, callback); } } const fetchContent = async (shopName,shopLink,callback) => { const { data } = await axios.get(shopLink); const $ = cheerio.load(data); let phone = $('.contact > p.phone').eq(0).text(); let email = $('.business-card-footer > a.email-business').eq(0).attr("href"); return callback(shopName,shopLink,phone,email); } async function scrapeData() { getLinks(startUrl,host,function(itemName,link){ fetchContent(itemName,link,function(shopName,shopLink,phone,email){ console.log({shopName,shopLink,phone,email}); }); }); } scrapeData();
补充说明
如果遇到抓取频率限制,可以在每次调用getLinks前加几秒的延时,避免被网站拦截。
内容的提问来源于stack exchange,提问作者robots.txt
相关产品推荐
相关产品推荐

