You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从HTML标签中提取纯文本?解决提取文本时的AttributeError报错问题

Fixing AttributeError and Extracting Clean Text from Indeed Job Listings

Hey there, let's get your code working properly. Here's what's going wrong and how to fix it:

What's Causing the Errors?

  • The AttributeError: When you use item.find(title='Python Developer'), some job listings don't have that exact title attribute, so this returns None. Trying to call get_text() on None throws the error you're seeing.
  • Unreliable Selection: Hardcoding the title attribute isn't a good approach—Indeed's job listings might have slight variations in titles, so we need a more consistent way to target job titles.

Corrected Code

import requests
from bs4 import BeautifulSoup

# Fetch the job page
response = requests.get("https://in.indeed.com/jobs?q=python%20developer&l=")
soup = BeautifulSoup(response.content, "html.parser")

# Locate the main results container (add a check in case it's missing)
results_body = soup.find(id="resultsBody")
if not results_body:
    print("Couldn't find the results section on the page.")
    exit()

# Get all individual job listing containers
job_containers = results_body.find_all(class_="slider_container")

for container in job_containers:
    # Target the actual job title element (this is the standard class on Indeed right now)
    title_element = container.find("a", class_="jcs-JobTitle")
    if title_element:
        # Extract clean, stripped text (removes extra spaces/newlines)
        job_title = title_element.get_text(strip=True)
        print(job_title)
    else:
        print("Skipping listing: Job title not found.")

Why This Works Better

  • Null Safety Checks: We now verify that both the results body and each job title element exist before trying to interact with them—no more AttributeError.
  • Consistent Element Targeting: Instead of chasing a specific title attribute, we use the <a> tag with class jcs-JobTitle (this is the standard class Indeed uses for job titles as of today). If the class ever changes, just inspect the page to find the new one.
  • Cleaner Text: get_text(strip=True) removes any extra whitespace or line breaks from the extracted text, giving you pure, readable titles.

Quick Notes

  • If you notice some listings are missing, Indeed loads some content dynamically with JavaScript. For those cases, you might need a tool like Selenium to render the full page.
  • Always make sure to follow Indeed's scraping guidelines to avoid getting blocked.

内容的提问来源于stack exchange,提问作者Ponaswin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 20:18:12