You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取avbuyer私人飞机数据不全&Wanted字段缺失求助

问题解决说明

核心问题根源

  • 数据缺失:你将数据写入temp列表的逻辑放在了list-other-dtl类div的遍历循环内,而Wanted、已售类的条目没有这个参数展示div,这类条目会直接被跳过,不会写入结果,这是你少了近50条数据的核心原因。
  • Price列无Wanted/Now Sold内容:原代码拿不到price类的元素时直接赋值为N/A,没有将对应状态内容同步到Price列。
  • 隐藏bug:updated字段的异常捕获分支错误给tag赋值了N/A,未给updated赋值,部分页面会触发变量未定义错误,中断当前条目处理。

修正后代码

import requests
from bs4 import BeautifulSoup
import pandas as pd
headers = {'User-Agent': 'Mozilla/5.0'}
temp=[]
# 可自行确认总页数后调整范围
for page in range(1, 20):
    response = requests.get(f"https://www.avbuyer.com/aircraft/private-jets/page-{page}", headers=headers)
    soup = BeautifulSoup(response.content, 'html.parser')
    postings = soup.find_all('div', class_='grid-x list-content')
    for post in postings:
        # 基础字段统一提取
        link = post.find('a').get('href')
        link_full = 'https://www.avbuyer.com' + link
        plane = post.find('h2', class_ = 'item-title').text.strip()
        # 调整价格提取逻辑,优先拿正常售价,无售价时填充Wanted/已售状态
        try:
            price = post.find('div', class_='price').text.strip()
        except:
            try:
                price = post.find('div', class_ = 'large-auto medium-auto cell').text.strip()
            except:
                price = "N/A"
        location = post.find('div', class_ = 'list-item-location').text.strip()
        desc = post.find('div', class_ = 'list-item-para').text.strip()
        try:
            tag = post.find('div', class_ = 'list-viewing-date').text.strip()
        except:
            tag = 'N/A'
        # 修复updated异常捕获的赋值错误
        try:
            updated = post.find('div', class_ = 'list-update').text.strip()
        except:
            updated = 'N/A'
        try:
            client = post.find('p', class_ = 'client-name').text.strip()
        except:
            client = 'N/A'
        try:
            highlight = post.find('span', class_ = 'red-text').text.strip()
        except:
            highlight = 'N/A'
        # 参数字段不存在时统一赋值N/A
        years = sn = total_time = 'N/A'
        t = post.find_all('div',class_='list-other-dtl')
        for i in t:
            data = [tup.text.strip() for tup in i.find_all('li')]
            if len(data) >=3:
                years = data[0]
                sn = data[1]
                total_time = data[2]
        # 所有条目统一写入,不再放在参数遍历的循环内
        temp.append([plane,price,location,years,sn,total_time,desc,tag,highlight,client,updated,link_full])

df=pd.DataFrame(temp,columns=["Plane","Price","location","Year","S/N","Total time","Description", "Tag", "Highlight", "Client", "Updated", "Link"])
# 编码选utf-8-sig避免打开乱码,同时去掉无用索引列
df.to_csv('/Users/xxx/avbuyer.csv', encoding='utf-8-sig', index=False)

调整说明

  • 把数据写入逻辑从参数遍历块中移出,所有条目无论是否有参数都会被写入结果
  • Price字段优先获取正常售价,拿不到时自动填充Wanted、Now Sold等状态值,无需额外单独存储状态列
  • 修复了updated字段的赋值bug,增加参数长度校验,避免参数不足时报错
  • 所有字段增加strip处理,清理多余的换行、空格
  • 导出CSV时增加编码声明,避免乱码,同时去掉无用的索引列

内容的提问来源于stack exchange,提问作者Alistair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 11:24:05