Python爬虫求助:属性缺失报错处理与CSV表头重复解决
Hey there! Let's fix those two annoying issues with your Immoscout24 crawler—they're totally solvable with a few targeted tweaks to your code. Let's break this down step by step:
Problem 1: Fixing NameError when attributes are missing
The root issue here is that if a container (like beschreibung_container) comes back empty, you never define the corresponding variable (e.g., beschreibung), so when you try to write it to CSV, Python throws an error.
The fix is to:
- Initialize all your output variables to empty strings upfront, so they always exist.
- Check if each container has content before trying to extract text from it.
- Use more specific error handling to catch issues without swallowing all mistakes.
Problem 2: Stopping repeated headers in CSV
Right now, you're writing the headers inside your loop over each property—so every time you process a new expose, you re-write the headers. We'll move the header write to before the loop, so it only runs once.
Modified Working Code
Here's your updated code with both fixes applied, plus some quality-of-life improvements:
from bs4 import BeautifulSoup import urllib.request as uReq from datetime import datetime # Generate filename with timestamp (kept your logic here) current_datetime = datetime.now().strftime("%Y%m%d_%H%M%S") filename = f"results_{current_datetime}.csv" # Write headers ONCE before processing any properties headers = "Objekt-ID##Titel##Adresse##Merkmale##Kosten##Bausubstanz und Energieausweis##Beschreibung##Ausstattung##Lage\n" with open(filename, "w", encoding="utf-8") as f: f.write(headers) # Assume 'numbers' is your list of expose IDs (keep your loop logic) for number in numbers: my_url = f"https://www.immobilienscout24.de/expose/{number}#/" uClient = uReq(my_url) page_html = uClient.read() uClient.close() page_soup = BeautifulSoup(page_html, "html.parser") containers = page_soup.find_all("div", {"id":"is24-content"}) # Initialize ALL fields to empty strings to avoid NameError objektid = "" titel = "" adresse = "" criteria = "" preis = "" energie = "" beschreibung = "" ausstattung = "" lage = "" for container in containers: try: # Extract Objekt-ID only if container exists objektid_container = container.find_all("div", {"class":"is24-scoutid__content padding-top-s"}) if objektid_container: objektid = objektid_container[0].get_text().strip() # Extract Titel titel_container = container.find_all("h1", {"class":"font-semibold font-xl margin-bottom margin-top-m palm-font-l"}) if titel_container: titel = titel_container[0].get_text().strip() # Extract Adresse adresse_container = container.find_all("div", {"class":"address-block"}) if adresse_container: adresse = adresse_container[0].get_text().strip() # Extract Merkmale (clean spaces here to keep code DRY) criteria_container = container.find_all("div", {"class":"criteriagroup criteria-group--two-columns"}) if criteria_container: criteria = criteria_container[0].get_text().strip().replace(" ", ";") # Extract Kosten preis_container = container.find_all("div", {"class":"grid-item lap-one-half desk-one-half padding-right-s"}) if preis_container: preis = preis_container[0].get_text().strip().replace(" ", ";") # Extract Bausubstanz und Energieausweis energie_container = container.find_all("div", {"class":"criteriagroup criteria-group--border criteria-group--two-columns criteria-group--spacing"}) if energie_container: energie = energie_container[0].get_text().strip().replace(" ", ";") # Extract Beschreibung beschreibung_container = container.find_all("pre", {"class":"is24qa-objektbeschreibung text-content short-text"}) if beschreibung_container: beschreibung = beschreibung_container[0].get_text().strip().replace("\n", " ") # Extract Ausstattung ausstattung_container = container.find_all("pre", {"class":"is24qa-ausstattung text-content short-text"}) if ausstattung_container: ausstattung = ausstattung_container[0].get_text().strip().replace("\n", " ") # Extract Lage lage_container = container.find_all("pre", {"class":"is24qa-lage text-content short-text"}) if lage_container: lage = lage_container[0].get_text().strip().replace("\n", " ") except Exception as e: # Print specific error + expose ID to debug easily print(f"Error processing expose {number}: {str(e)}") pass # Write the row to CSV (even if some fields are empty) row = f"{objektid}##{titel}##{adresse}##{criteria}##{preis}##{energie}##{beschreibung}##{ausstattung}##{lage}\n" with open(filename, "a", encoding="utf-8") as f: f.write(row)
Key Changes Explained
- Empty String Initialization: Every field starts as
"", so even if a container is missing, the variable still exists—no moreNameError. - Container Checks: For each field, we first check if the
find_allresult is non-empty before accessing[0]—preventsIndexErrorwhen elements are missing. - Single Header Write: Headers are written once before the loop starts, using
"w"mode to create the file. - Append Mode for Rows: Each property row is added with
"a"mode, which appends to the file instead of overwriting it. - UTF-8 Encoding: Added
encoding="utf-8"to handle German special characters (ä, ö, ü) without garbled text. - Better Error Logging: Prints the expose ID and specific error message, so you know exactly which property caused a problem.
内容的提问来源于stack exchange,提问作者Juli Goe

