Python爬虫遇AttributeError,如何提取Indeed有效公司名称?
Fix AttributeError & Show Valid Company Names for Indeed AU Scraper
Let's break down the issues with your code and fix them step by step:
1. The Root Cause of the AttributeError
Your logic for checking the company variable is completely reversed:
- When
companyisNone, you’re trying to callcompany.find("a")— which is impossible becauseNonedoesn’t have afindmethod. - You also didn’t account for cases where company names are nested inside an
<a>tag within thecompanyspan, leading to empty/Noneoutputs.
2. Fixed Code with Explanations
Here’s the revised version of your scraper, with robust error handling and logic fixes:
import requests from bs4 import BeautifulSoup LIMIT = 50 URL = f"https://au.indeed.com/jobs?q=Python&limit={LIMIT}&radius=50" def extract_indeed_pages(): result = requests.get(URL) soup = BeautifulSoup(result.text, "html.parser") pagination = soup.find("div", {"class":"pagination"}) # Handle edge case where no pagination exists (only 1 page) if not pagination: return 1 links = pagination.find_all("a") pages = [] for link in links[:-1]: # Ensure we only process valid page numbers page_text = link.string if page_text: pages.append(int(page_text)) max_page = pages[-1] if pages else 1 return max_page def extract_indeed_jobs(last_page): jobs = [] # Uncomment this loop if you want to scrape all pages instead of just the first # for page in range(last_page): result = requests.get(f"{URL}&start={0*LIMIT}") soup = BeautifulSoup(result.text, "html.parser") results = soup.find_all("div", {"class": "jobsearch-SerpJobCard"}) for result in results: # Extract job title with fallback title_elem = result.find("div", {"class": "title"}) title = title_elem.find("a")["title"] if title_elem else "Unknown Job Title" # Extract company name with proper validation company = result.find("span", {"class": "company"}) company_name = None if company: # Check if company name is in an <a> tag (common for sponsored jobs) company_link = company.find("a") if company_link: company_name = company_link.string.strip() else: # Company name is directly in the span company_name = company.string.strip() # Only print and save valid company names if company_name: print(company_name) jobs.append({ "title": title, "company": company_name }) else: # Optional: mark jobs with no valid company name print("Unknown Company") return jobs # Run the scraper last_indeed_page = extract_indeed_pages() indeed_jobs = extract_indeed_jobs(last_indeed_page)
3. Key Fixes Made
- Corrected company check logic: We now verify
companyexists before trying to access its contents, eliminating theAttributeError. - Handled both company name structures: Indeed uses two formats for company names — directly in the
<span>or nested in an<a>tag — we cover both cases. - Cleaned up text: Added
.strip()to remove extra whitespace from company names. - Filtered invalid entries: Only non-empty, valid company names are printed and added to the jobs list, getting rid of those
Noneoutputs. - Robust pagination: Added checks to handle cases where there’s only one page of results (no pagination bar).
4. Why This Works
We never call methods like find() or access string on a None object anymore — every element is validated first. We also explicitly account for Indeed’s HTML structure variations, ensuring we always pull valid company names when they exist.
内容的提问来源于stack exchange,提问作者조조명현
相关产品推荐
相关产品推荐

