You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫遇AttributeError,如何提取Indeed有效公司名称?

Fix AttributeError & Show Valid Company Names for Indeed AU Scraper

Let's break down the issues with your code and fix them step by step:

1. The Root Cause of the AttributeError

Your logic for checking the company variable is completely reversed:

  • When company is None, you’re trying to call company.find("a") — which is impossible because None doesn’t have a find method.
  • You also didn’t account for cases where company names are nested inside an <a> tag within the company span, leading to empty/None outputs.

2. Fixed Code with Explanations

Here’s the revised version of your scraper, with robust error handling and logic fixes:

import requests
from bs4 import BeautifulSoup
LIMIT = 50
URL = f"https://au.indeed.com/jobs?q=Python&limit={LIMIT}&radius=50"

def extract_indeed_pages():
    result = requests.get(URL)
    soup = BeautifulSoup(result.text, "html.parser")
    pagination = soup.find("div", {"class":"pagination"})
    # Handle edge case where no pagination exists (only 1 page)
    if not pagination:
        return 1
    links = pagination.find_all("a")
    pages = []
    for link in links[:-1]:
        # Ensure we only process valid page numbers
        page_text = link.string
        if page_text:
            pages.append(int(page_text))
    max_page = pages[-1] if pages else 1
    return max_page

def extract_indeed_jobs(last_page):
    jobs = []
    # Uncomment this loop if you want to scrape all pages instead of just the first
    # for page in range(last_page):
    result = requests.get(f"{URL}&start={0*LIMIT}")
    soup = BeautifulSoup(result.text, "html.parser")
    results = soup.find_all("div", {"class": "jobsearch-SerpJobCard"})
    for result in results:
        # Extract job title with fallback
        title_elem = result.find("div", {"class": "title"})
        title = title_elem.find("a")["title"] if title_elem else "Unknown Job Title"
        
        # Extract company name with proper validation
        company = result.find("span", {"class": "company"})
        company_name = None
        if company:
            # Check if company name is in an <a> tag (common for sponsored jobs)
            company_link = company.find("a")
            if company_link:
                company_name = company_link.string.strip()
            else:
                # Company name is directly in the span
                company_name = company.string.strip()
        
        # Only print and save valid company names
        if company_name:
            print(company_name)
            jobs.append({
                "title": title,
                "company": company_name
            })
        else:
            # Optional: mark jobs with no valid company name
            print("Unknown Company")
    return jobs

# Run the scraper
last_indeed_page = extract_indeed_pages()
indeed_jobs = extract_indeed_jobs(last_indeed_page)

3. Key Fixes Made

  • Corrected company check logic: We now verify company exists before trying to access its contents, eliminating the AttributeError.
  • Handled both company name structures: Indeed uses two formats for company names — directly in the <span> or nested in an <a> tag — we cover both cases.
  • Cleaned up text: Added .strip() to remove extra whitespace from company names.
  • Filtered invalid entries: Only non-empty, valid company names are printed and added to the jobs list, getting rid of those None outputs.
  • Robust pagination: Added checks to handle cases where there’s only one page of results (no pagination bar).

4. Why This Works

We never call methods like find() or access string on a None object anymore — every element is validated first. We also explicitly account for Indeed’s HTML structure variations, ensuring we always pull valid company names when they exist.

内容的提问来源于stack exchange,提问作者조조명현

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:06:51