使用Cheerio+Node.js爬取MonsterIndia零经验PHP职位的问题
Hey there! As someone new to web scraping with Node.js and Cheerio, it makes total sense to get stuck when filter parameters don't show up in the URL—dynamic sites like MonsterIndia love hiding their data behind AJAX requests instead of updating the URL directly. Let's break down how to fix this:
Step 1: Understand how the site handles filters
When you click the "Zero Experience" filter and the URL stays unchanged, that means the site loads filtered jobs dynamically via JavaScript (usually an AJAX call to an internal API endpoint). To uncover what's happening, open your browser's DevTools (F12), switch to the Network tab, then click the filter again. You'll see a new request pop up—this is the one fetching the filtered job data.
Step 2: Extract the API details
Examine that new request closely:
- Note the request URL (it might look like
https://www.monsterindia.com/job-search-api/v3/search) - Check the request method (most likely POST, but could be GET)
- Inspect the request body or query parameters—you'll find a field tied to experience, probably something like
experience: "0"orexperienceRange: "0-0" - Don't forget to copy key request headers like
User-Agent(to mimic a real browser) and any required cookies
Step 3: Adjust your code to target the API
Instead of requesting the initial PHP jobs page, send a request directly to that API endpoint with the correct parameters. Here's how to modify your existing request-based code:
var request = require('request'); var cheerio = require('cheerio'); // You may not need this if the API returns JSON! // Configure options based on what you found in DevTools const scrapeOptions = { url: 'https://www.monsterindia.com/job-search-api/v3/search', method: 'POST', headers: { 'Content-Type': 'application/json', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', // Add any required cookies or auth headers if observed in DevTools }, body: JSON.stringify({ query: 'php', experience: '0', // This is the zero-experience filter parameter location: '', // Leave empty for all locations, or add a city if needed page: 1, sort: 'relevance' }) }; request(scrapeOptions, function(error, response, body) { if (!error && response.statusCode === 200) { // Most job APIs return JSON, so parse it directly const jobData = JSON.parse(body); // Loop through results and extract key details jobData.data.forEach(job => { console.log(`Job Title: ${job.title}`); console.log(`Company: ${job.companyName}`); console.log(`Experience Required: ${job.experience}`); console.log(`Job URL: https://www.monsterindia.com/job-openings/${job.jobId}`); console.log('---'); }); } else { console.error('Something went wrong:', error || `Status Code: ${response.statusCode}`); } });
Key Tips to Avoid Headaches
- Verify API details regularly: MonsterIndia might update their API endpoints or parameters, so double-check with DevTools if your code stops working.
- Bypass anti-scraping measures: Sites often block requests that don't look like real browsers. Include a valid
User-Agentheader, and if you hit blocks, try adding cookies from a real browser session. - Cheerio vs. JSON: If the API returns HTML instead of JSON, you can still use Cheerio to parse the response body—just pass
bodytocheerio.load()like you would with a regular page. - Alternative: Use Puppeteer for click simulation: If you want to mimic the exact user flow (clicking the filter button), use Puppeteer (a headless browser) instead of
request—though this is slower than hitting the API directly.
内容的提问来源于stack exchange,提问作者Aashish Kumar

