You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup find_previous_sibling爬取车辆年份/公里数返回N/A问题

问题原因与解决方案

核心问题分析

你的代码爬取年份(Año)和公里数(Kilómetros)时失效,主要有两个关键原因:

  • 文本匹配不严谨:直接用text='Año'查找标题时,页面HTML中的文本可能包含首尾空格或换行符,导致精确匹配失败。
  • 兄弟节点查找逻辑错误:页面里每个属性(年份、公里数等)都被单独的div.carone-car-attribute容器包裹,值和标题是该容器内的子元素,直接跨容器用find_previous_sibling查找会出错;燃油类型能正常获取只是巧合,其DOM结构刚好符合你的错误查找逻辑。

修正后的代码实现

把属性提取逻辑改成遍历每个独立属性容器,精准匹配标题并提取对应值:

import pandas as pd
from datetime import date
import os
import socket
import requests
from bs4 import BeautifulSoup

def scrape_product_data(url):
    try:
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/58.0.3029.110 Safari/537.3"
        }

        product_data = []

        response = requests.get(url, headers=headers)
        response.raise_for_status()

        soup = BeautifulSoup(response.text, 'html.parser')
        product_elements = soup.find_all('div', class_='product-item-info')
        for product_element in product_elements:
            # 提取基础信息(保留原有逻辑)
            product_name_element = product_element.select_one('p.carone-car-info-data-brand.cursor-pointer')
            product_name = product_name_element.text.strip() if product_name_element else "N/A"

            product_price_element = product_element.find('span', class_='price')
            product_price = product_price_element.text.strip() if product_price_element else "N/A"

            product_model_element = product_element.select_one('p.carone-car-info-data-model')
            product_model = product_model_element.get('title').strip() if product_model_element else "N/A"

            # 重新实现属性提取逻辑
            attributes_div = product_element.find('div', class_='carone-car-attributes')
            # 初始化默认值
            year_value = "N/A"
            kilometers_value = "N/A"
            fuel_value = "N/A"

            if attributes_div:
                # 遍历所有独立属性容器
                attribute_items = attributes_div.find_all('div', class_='carone-car-attribute')
                for item in attribute_items:
                    title = item.find('p', class_='carone-car-attribute-title')
                    value = item.find('p', class_='carone-car-attribute-value')
                    if title and value:
                        clean_title = title.text.strip()
                        clean_value = value.text.strip()
                        if clean_title == 'Año':
                            year_value = clean_value
                        elif clean_title == 'Kilómetros':
                            kilometers_value = clean_value
                        elif clean_title == 'Combustible':
                            fuel_value = clean_value

            product_data.append((product_name, product_price, product_model, year_value, kilometers_value, fuel_value))
        
        return product_data
    except Exception as e:
        print(f"爬取出错: {str(e)}")
        return []

# 测试调用示例
scraped_data = scrape_product_data("https://carone.com.uy/autos-usados-y-0km?p=21")
for entry in scraped_data:
    print(entry)

关键修改点说明

  1. 遍历独立属性容器:通过find_all('div', class_='carone-car-attribute')获取每个属性的独立容器,确保标题和值在同一范围内处理,避免跨容器查找错误。
  2. 文本清洗匹配:对标题文本做strip()处理后再匹配,彻底避免空格、换行符导致的匹配失败。
  3. 初始化默认值:提前给所有属性赋值默认的"N/A",避免因找不到元素引发的报错。

内容的提问来源于stack exchange,提问作者Bruno

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:36:31