Next.js 14 Server Actions中用Cheerio无法爬取Beatport艺术家数据
问题分析与解决方案
问题根源
- 静态抓取无法获取动态渲染内容:Beatport的搜索结果大概率是通过客户端JavaScript动态加载的。你用
fetch拿到的是服务器返回的初始静态HTML,此时艺术家卡片还未被JS渲染出来——浏览器开发者工具看到的是渲染后的DOM结构,而Cheerio解析的是原始响应,自然找不到对应元素。 - 依赖不稳定的动态类名:
ArtistCard-style__Wrapper-sc-7ba2494f-10这类类名是前端构建工具(如CSS Modules)生成的哈希类名,网站更新时会随时变化,完全不可靠。 - 请求头未模拟浏览器:服务器可能通过检查
User-Agent等请求头识别出爬虫请求,返回不含完整内容的简化页面。
解决步骤
1. 替换为稳定的选择器
放弃动态生成的类名,改用页面中固定的测试属性或稳定结构:
- 页面内的艺术家卡片带有
data-testid="artist-card"属性,这是专门给自动化测试用的,不会轻易变更,用它定位卡片更可靠。 - 提取链接时,无需依赖
title属性精准匹配,直接查找卡片内的.artwork类链接即可。
2. 模拟浏览器请求头
给fetch添加浏览器风格的请求头,避免被识别为爬虫:
const headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8' }; const searchResponse = await fetch(searchUrl, { headers });
3. 处理动态内容(核心解决方法)
如果添加请求头后仍无法获取内容,说明内容是JS动态加载的,此时需要用支持JavaScript渲染的工具(如playwright-core),在Next.js Server Action中配置如下:
首先安装依赖:
npm install playwright-core
修改后的代码:
"use server"; import { chromium } from 'playwright-core'; import { load } from "cheerio"; interface BeatportArtist { name: string; beatportUrl: string; imageUrl: string; } const BASE_URL = "https://www.beatport.com"; export async function scrapeBeatportArtist( name: string ): Promise<BeatportArtist | null> { let browser; try { const searchUrl = `${BASE_URL}/search?q=${encodeURIComponent(name)}`; console.log(`Searching Beatport for artist: ${name}`); // 启动无头浏览器 browser = await chromium.launch({ headless: true }); const page = await browser.newPage(); await page.setExtraHTTPHeaders({ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' }); // 等待页面加载完成,直到艺术家卡片出现 await page.goto(searchUrl, { waitUntil: 'networkidle' }); await page.waitForSelector('[data-testid="artist-card"]', { timeout: 5000 }); // 获取渲染后的完整HTML const searchHtml = await page.content(); const $search = load(searchHtml); // 用稳定的data-testid定位第一个艺术家卡片 const artistCard = $search('[data-testid="artist-card"]').first(); if (!artistCard.length) { console.log(`No artist card found for artist: ${name}`); return null; } // 提取艺术家主页链接 const artistLink = artistCard.find('a.artwork').attr("href"); if (!artistLink) { console.log(`No Beatport profile found for artist: ${name}`); return null; } const artistUrl = `${BASE_URL}${artistLink}`; // 爬取艺术家详情页 await page.goto(artistUrl, { waitUntil: 'networkidle' }); const artistHtml = await page.content(); const $artist = load(artistHtml); const imageUrl = $artist(".artist-hero__image img").attr("src") || ""; return { name, beatportUrl: artistUrl, imageUrl, }; } catch (error) { console.error(`Error scraping Beatport for artist ${name}:`, error); return null; } finally { if (browser) await browser.close(); } }
备选方案:直接调用隐藏API
打开浏览器开发者工具的Network标签,搜索艺术家时观察XHR/fetch请求,Beatport可能会有一个API接口返回搜索结果的JSON数据(例如类似https://api.beatport.com/v4/search?q=xxx的端点)。直接请求这个API比爬取页面更高效可靠,无需处理DOM解析。
验证要点
- 先用
console.log(searchHtml)输出fetch到的原始HTML,检查是否包含艺术家卡片内容,以此确认是否为动态加载问题。 - 在浏览器控制台用
document.querySelector('[data-testid="artist-card"]')验证选择器是否有效。
内容的提问来源于stack exchange,提问作者kelvin
相关产品推荐
相关产品推荐

