You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping及BeautifulSoup下一页解析与航空数据CSV导出问题

问题原因

现有代码仅完成了第一页的数据提取、第二页的页面请求,没有对第二页及后续页面的列表数据做提取逻辑,也没有循环遍历所有分页,最终CSV仅存储第一页数据。另外存在两个可优化点:

  • 定义了headers请求头但未传入requests.get方法,容易触发站点反爬策略
  • 两次页面解析使用了不同的解析器(html.parser和lxml),可能出现解析结果不一致的问题
修复后完整代码
import requests
from bs4 import BeautifulSoup
import pandas as pd
import time

headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}
base_url = 'https://www.avbuyer.com'
current_url = 'https://www.avbuyer.com/aircraft/private-jets'
temp = []

while True:
    # 请求当前页
    response = requests.get(current_url, headers=headers)
    soup = BeautifulSoup(response.content, 'lxml')
    
    # 提取当前页所有列表数据
    postings = soup.find_all('div', class_ = 'listing-item premium')
    for post in postings:
        link = post.find('a', class_ = 'more-info').get('href')
        link_full = base_url + link
        plane = post.find('h2', class_ = 'item-title').text.strip()
        price = post.find('div', class_ = 'price').text.strip()
        location = post.find('div', class_ = 'list-item-location').text.strip()
        desc = post.find('div', class_ = 'list-item-para').text.strip()
        try:
            tag = post.find('div', class_ = 'list-viewing-date').text.strip()
        except:
            tag = 'N/A'
        updated = post.find('div', class_ = 'list-update').text.strip()
        t = post.find_all('div',class_='list-other-dtl')
        for i in t:
            data = [tup.text.strip() for tup in i.find_all('li')]
            years = data[0]
            s = data[1]
            total_time = data[2]
            temp.append([plane,price,location,years,s,total_time,desc,tag,updated,link_full])
    
    # 查找下一页链接,没有则退出循环
    next_tag = soup.find('a', {'rel':'next'})
    if not next_tag:
        break
    next_page = next_tag.get('href')
    current_url = base_url + next_page
    # 加延时避免被封
    time.sleep(2)

# 导出CSV
df = pd.DataFrame(temp, columns=["plane","price","location","Year","S/N","Totaltime","Description","Tag","Last Updated","link"])
df.to_csv('/Users/xxx/avbuyer.csv', index=False, encoding='utf-8-sig')
注意事项
  • 代码中加入了time.sleep(2)的请求间隔,避免请求频率过高被站点封禁IP,可根据实际情况调整间隔时长
  • 导出CSV时指定了encoding='utf-8-sig',避免Windows系统下打开CSV出现乱码
  • 如果需要爬取非置顶的普通私人机信息,可以将列表选择器从listing-item premium调整为匹配所有listing-item类的div
  • 运行环境不会影响执行结果,Python 3.8.12版本完全兼容所有依赖库,Chrome浏览器未参与爬虫请求流程,和运行结果无关

内容的提问来源于stack exchange,提问作者Alistair

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 05:36:03