BeautifulSoup find_previous_sibling爬取车辆年份/公里数返回N/A问题
问题原因与解决方案
核心问题分析
你的代码爬取年份(Año)和公里数(Kilómetros)时失效,主要有两个关键原因:
- 文本匹配不严谨:直接用
text='Año'查找标题时,页面HTML中的文本可能包含首尾空格或换行符,导致精确匹配失败。 - 兄弟节点查找逻辑错误:页面里每个属性(年份、公里数等)都被单独的
div.carone-car-attribute容器包裹,值和标题是该容器内的子元素,直接跨容器用find_previous_sibling查找会出错;燃油类型能正常获取只是巧合,其DOM结构刚好符合你的错误查找逻辑。
修正后的代码实现
把属性提取逻辑改成遍历每个独立属性容器,精准匹配标题并提取对应值:
import pandas as pd from datetime import date import os import socket import requests from bs4 import BeautifulSoup def scrape_product_data(url): try: headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3" } product_data = [] response = requests.get(url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') product_elements = soup.find_all('div', class_='product-item-info') for product_element in product_elements: # 提取基础信息(保留原有逻辑) product_name_element = product_element.select_one('p.carone-car-info-data-brand.cursor-pointer') product_name = product_name_element.text.strip() if product_name_element else "N/A" product_price_element = product_element.find('span', class_='price') product_price = product_price_element.text.strip() if product_price_element else "N/A" product_model_element = product_element.select_one('p.carone-car-info-data-model') product_model = product_model_element.get('title').strip() if product_model_element else "N/A" # 重新实现属性提取逻辑 attributes_div = product_element.find('div', class_='carone-car-attributes') # 初始化默认值 year_value = "N/A" kilometers_value = "N/A" fuel_value = "N/A" if attributes_div: # 遍历所有独立属性容器 attribute_items = attributes_div.find_all('div', class_='carone-car-attribute') for item in attribute_items: title = item.find('p', class_='carone-car-attribute-title') value = item.find('p', class_='carone-car-attribute-value') if title and value: clean_title = title.text.strip() clean_value = value.text.strip() if clean_title == 'Año': year_value = clean_value elif clean_title == 'Kilómetros': kilometers_value = clean_value elif clean_title == 'Combustible': fuel_value = clean_value product_data.append((product_name, product_price, product_model, year_value, kilometers_value, fuel_value)) return product_data except Exception as e: print(f"爬取出错: {str(e)}") return [] # 测试调用示例 scraped_data = scrape_product_data("https://carone.com.uy/autos-usados-y-0km?p=21") for entry in scraped_data: print(entry)
关键修改点说明
- 遍历独立属性容器:通过
find_all('div', class_='carone-car-attribute')获取每个属性的独立容器,确保标题和值在同一范围内处理,避免跨容器查找错误。 - 文本清洗匹配:对标题文本做
strip()处理后再匹配,彻底避免空格、换行符导致的匹配失败。 - 初始化默认值:提前给所有属性赋值默认的"N/A",避免因找不到元素引发的报错。
内容的提问来源于stack exchange,提问作者Bruno
相关产品推荐
相关产品推荐

