You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Cheerio+Node.js爬取MonsterIndia零经验PHP职位的问题

解决MonsterIndia零经验PHP岗位爬取问题

Hey there! As someone new to web scraping with Node.js and Cheerio, it makes total sense to get stuck when filter parameters don't show up in the URL—dynamic sites like MonsterIndia love hiding their data behind AJAX requests instead of updating the URL directly. Let's break down how to fix this:

Step 1: Understand how the site handles filters

When you click the "Zero Experience" filter and the URL stays unchanged, that means the site loads filtered jobs dynamically via JavaScript (usually an AJAX call to an internal API endpoint). To uncover what's happening, open your browser's DevTools (F12), switch to the Network tab, then click the filter again. You'll see a new request pop up—this is the one fetching the filtered job data.

Step 2: Extract the API details

Examine that new request closely:

  • Note the request URL (it might look like https://www.monsterindia.com/job-search-api/v3/search)
  • Check the request method (most likely POST, but could be GET)
  • Inspect the request body or query parameters—you'll find a field tied to experience, probably something like experience: "0" or experienceRange: "0-0"
  • Don't forget to copy key request headers like User-Agent (to mimic a real browser) and any required cookies

Step 3: Adjust your code to target the API

Instead of requesting the initial PHP jobs page, send a request directly to that API endpoint with the correct parameters. Here's how to modify your existing request-based code:

var request = require('request');
var cheerio = require('cheerio'); // You may not need this if the API returns JSON!

// Configure options based on what you found in DevTools
const scrapeOptions = {
  url: 'https://www.monsterindia.com/job-search-api/v3/search',
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    // Add any required cookies or auth headers if observed in DevTools
  },
  body: JSON.stringify({
    query: 'php',
    experience: '0', // This is the zero-experience filter parameter
    location: '', // Leave empty for all locations, or add a city if needed
    page: 1,
    sort: 'relevance'
  })
};

request(scrapeOptions, function(error, response, body) {
  if (!error && response.statusCode === 200) {
    // Most job APIs return JSON, so parse it directly
    const jobData = JSON.parse(body);
    
    // Loop through results and extract key details
    jobData.data.forEach(job => {
      console.log(`Job Title: ${job.title}`);
      console.log(`Company: ${job.companyName}`);
      console.log(`Experience Required: ${job.experience}`);
      console.log(`Job URL: https://www.monsterindia.com/job-openings/${job.jobId}`);
      console.log('---');
    });
  } else {
    console.error('Something went wrong:', error || `Status Code: ${response.statusCode}`);
  }
});

Key Tips to Avoid Headaches

  • Verify API details regularly: MonsterIndia might update their API endpoints or parameters, so double-check with DevTools if your code stops working.
  • Bypass anti-scraping measures: Sites often block requests that don't look like real browsers. Include a valid User-Agent header, and if you hit blocks, try adding cookies from a real browser session.
  • Cheerio vs. JSON: If the API returns HTML instead of JSON, you can still use Cheerio to parse the response body—just pass body to cheerio.load() like you would with a regular page.
  • Alternative: Use Puppeteer for click simulation: If you want to mimic the exact user flow (clicking the filter button), use Puppeteer (a headless browser) instead of request—though this is slower than hitting the API directly.

内容的提问来源于stack exchange,提问作者Aashish Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:11:02