You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy从Indeed职位列表提取薪资信息时遇到提取失败问题

使用Scrapy从Indeed职位列表提取薪资信息时遇到提取失败问题

Hey there! Let's troubleshoot why you're not getting the salary data to show up correctly.

First off, I notice in your code you're grabbing the salarySnippet directly from the job object, but your output shows "N/A" even though you've seen the salary data exists in the JSON response. Here's what's likely going on and how to fix it:

1. Verify the exact key path in the JSON structure

Sometimes job listings without salary info won't have the salarySnippet key at all, which is why your get() method falls back to the default {}. But if you've confirmed some jobs do have this data, let's make sure you're accessing the right spot.

Add a quick line to print all keys in the job object so you can confirm where the salary data lives:

print(job.keys())

Or even better, print the full job JSON to inspect the structure clearly:

import json
print(json.dumps(job, indent=2))

This will show you if salarySnippet is directly under the job object, or nested under another key like compensation or salaryMetadata.

2. Extract specific fields from salarySnippet

Once you confirm salarySnippet is the correct key, you need to pull the actual text and currency values from that dictionary instead of printing the whole thing. Modify your salary extraction code like this:

# Extract salary details from salarySnippet
salary_snippet = job.get('salarySnippet', {})
salary_text = salary_snippet.get('text', 'N/A')
salary_currency = salary_snippet.get('currency', 'N/A')

Then update your print statement to show the readable salary info:

print(f"Job Name: {job_name}")
print(f"Job Location: {job_location}")
print(f"Salary: {salary_text} ({salary_currency})")
print("-" * 40)

This way, if the salary data exists, you'll get something like ₹40,000 - ₹70,000 a month (INR); if not, it'll show "N/A" for both fields gracefully.

3. Handle edge cases

Keep in mind that not all Indeed job listings include salary information, so seeing "N/A" for some entries is totally normal. Your code already handles this well with the get() method's default values.

Putting it all together, your updated parse method would look like this:

def parse(self, response):
    # Extract job titles
    script_tag = re.findall(r'window.mosaic.providerData\["mosaic-provider-jobcards"\]=(\{.+?\});', response.text)
    if script_tag:
        json_blob = json.loads(script_tag[0])
        jobs_list = json_blob['metaData']['mosaicProviderJobCardsModel']['results']
        for job in jobs_list:
            job_name = job.get('displayTitle', 'N/A')
            job_location = job.get('formattedLocation', 'N/A')
            
            # Extract salary details
            salary_snippet = job.get('salarySnippet', {})
            salary_text = salary_snippet.get('text', 'N/A')
            salary_currency = salary_snippet.get('currency', 'N/A')
            
            print(f"Job Name: {job_name}")
            print(f"Job Location: {job_location}")
            print(f"Salary: {salary_text} ({salary_currency})")
            print("-" * 40)

Give this a try—you should start seeing the salary data for listings that have it!

备注:内容来源于stack exchange,提问作者Mhd jaseer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 15:12:59