如何从HTML标签中提取纯文本?解决提取文本时的AttributeError报错问题
Fixing AttributeError and Extracting Clean Text from Indeed Job Listings
Hey there, let's get your code working properly. Here's what's going wrong and how to fix it:
What's Causing the Errors?
- The AttributeError: When you use
item.find(title='Python Developer'), some job listings don't have that exact title attribute, so this returnsNone. Trying to callget_text()onNonethrows the error you're seeing. - Unreliable Selection: Hardcoding the title attribute isn't a good approach—Indeed's job listings might have slight variations in titles, so we need a more consistent way to target job titles.
Corrected Code
import requests from bs4 import BeautifulSoup # Fetch the job page response = requests.get("https://in.indeed.com/jobs?q=python%20developer&l=") soup = BeautifulSoup(response.content, "html.parser") # Locate the main results container (add a check in case it's missing) results_body = soup.find(id="resultsBody") if not results_body: print("Couldn't find the results section on the page.") exit() # Get all individual job listing containers job_containers = results_body.find_all(class_="slider_container") for container in job_containers: # Target the actual job title element (this is the standard class on Indeed right now) title_element = container.find("a", class_="jcs-JobTitle") if title_element: # Extract clean, stripped text (removes extra spaces/newlines) job_title = title_element.get_text(strip=True) print(job_title) else: print("Skipping listing: Job title not found.")
Why This Works Better
- Null Safety Checks: We now verify that both the results body and each job title element exist before trying to interact with them—no more AttributeError.
- Consistent Element Targeting: Instead of chasing a specific title attribute, we use the
<a>tag with classjcs-JobTitle(this is the standard class Indeed uses for job titles as of today). If the class ever changes, just inspect the page to find the new one. - Cleaner Text:
get_text(strip=True)removes any extra whitespace or line breaks from the extracted text, giving you pure, readable titles.
Quick Notes
- If you notice some listings are missing, Indeed loads some content dynamically with JavaScript. For those cases, you might need a tool like Selenium to render the full page.
- Always make sure to follow Indeed's scraping guidelines to avoid getting blocked.
内容的提问来源于stack exchange,提问作者Ponaswin
相关产品推荐
相关产品推荐

