求助:使用Scrapy爬取Indeed职位页面无法获取数据
Let's figure out why your Scrapy spider isn't pulling in the job description and company name data, and get it fixed up step by step.
1. Outdated Selectors (Most Likely Culprit)
Indeed regularly updates its page structure, so the XPath selectors you're using might no longer match the current elements. Let's adjust them to match the latest Indeed UK layout:
Company Name Selector
Your current selector //*[contains(@class, "icl-u-lg-mr--sm")]//text() targets a class that's likely been retired. On modern Indeed job pages, the company name is consistently found in an element with the data-testid="company-name" attribute. Replace your company selector with:
company = response.xpath('//div[@data-testid="company-name"]/text()').extract_first()
Job Description Selector
While //*[@id="jobDescriptionText"] might still work, wrapping the text extraction with a cleaner join will ensure you capture all content without extra line breaks. Update this line to:
job_description = " ".join(response.xpath('//div[@id="jobDescriptionText"]//text()').extract()).strip()
2. Dynamic Content Blocking
Some parts of Indeed's job pages (especially longer descriptions) load dynamically via JavaScript. Scrapy's default Request doesn't execute JS, so you might be getting a partial or empty HTML response. To fix this, use scrapy-playwright to render pages like a real browser:
- First install the package:
pip install scrapy-playwright - Add these settings to your
settings.py:DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, "args": ["--no-sandbox"], } - Modify your job page request to enable Playwright:
yield Request( url=absulate_job_link, callback=self.parse_jobpage, meta={ "Job Title": job_title, "Location": job_location, "Job Link": absulate_job_link, "playwright": True } )
3. Anti-Scraping Measures
Indeed actively blocks scrapers. Mimic a real user to avoid being detected:
- Add a realistic user-agent to
settings.py:USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36" - Slow down requests to avoid triggering rate limits:
DOWNLOAD_DELAY = 2 # Wait 2 seconds between requests
4. Verify What Scrapy Is Receiving
Before tweaking selectors, always check the raw response Scrapy gets. Add this line at the start of parse_jobpage to save the page to a file:
with open("job_page.html", "wb") as f: f.write(response.body)
Open the saved HTML in a browser—if the job description/company name are missing here, it's either dynamic content or anti-scraping blocking you.
Updated Working Spider Code
Here's your spider with all the fixes applied:
from scrapy import Spider from scrapy.http import Request class IndeedSpider(Spider): name = 'indeed' allowed_domains = ['indeed.com', 'indeed.co.uk'] start_urls = ['https://www.indeed.co.uk/jobs?q=Russian&fromage=1'] def parse(self, response): # Updated job row selector for current Indeed layout jobs = response.xpath('//div[contains(@class, "css-5lfssm eu4oa1w0")]') for job in jobs: job_title = job.xpath('.//h2[@class="jobTitle css-1h4a4n5 eu4oa1w0"]//a/@title').extract_first() job_location = job.xpath('.//div[@class="css-1p0sjhy eu4oa1w0"]/text()').extract_first() job_link = job.xpath('.//h2[@class="jobTitle css-1h4a4n5 eu4oa1w0"]//a/@href').extract_first() if job_link: absulate_job_link = response.urljoin(job_link) self.logger.info(f"Scraping job: {absulate_job_link}") yield Request( url=absulate_job_link, callback=self.parse_jobpage, meta={ "Job Title": job_title, "Location": job_location, "Job Link": absulate_job_link } ) def parse_jobpage(self, response): job_title = response.meta.get('Job Title') job_location = response.meta.get('Location') absulate_job_link = response.meta.get('Job Link') job_description = " ".join(response.xpath('//div[@id="jobDescriptionText"]//text()').extract()).strip() company = response.xpath('//div[@data-testid="company-name"]/text()').extract_first() yield { "Job Title": job_title, "Location": job_location, "Job Link": absulate_job_link, "Job Description": job_description, "Company": company }
内容的提问来源于stack exchange,提问作者booleantrue

