使用BeautifulSoup爬取avbuyer私人飞机数据不全&Wanted字段缺失求助
问题解决说明
核心问题根源
- 数据缺失:你将数据写入
temp列表的逻辑放在了list-other-dtl类div的遍历循环内,而Wanted、已售类的条目没有这个参数展示div,这类条目会直接被跳过,不会写入结果,这是你少了近50条数据的核心原因。 - Price列无Wanted/Now Sold内容:原代码拿不到
price类的元素时直接赋值为N/A,没有将对应状态内容同步到Price列。 - 隐藏bug:
updated字段的异常捕获分支错误给tag赋值了N/A,未给updated赋值,部分页面会触发变量未定义错误,中断当前条目处理。
修正后代码
import requests from bs4 import BeautifulSoup import pandas as pd headers = {'User-Agent': 'Mozilla/5.0'} temp=[] # 可自行确认总页数后调整范围 for page in range(1, 20): response = requests.get(f"https://www.avbuyer.com/aircraft/private-jets/page-{page}", headers=headers) soup = BeautifulSoup(response.content, 'html.parser') postings = soup.find_all('div', class_='grid-x list-content') for post in postings: # 基础字段统一提取 link = post.find('a').get('href') link_full = 'https://www.avbuyer.com' + link plane = post.find('h2', class_ = 'item-title').text.strip() # 调整价格提取逻辑,优先拿正常售价,无售价时填充Wanted/已售状态 try: price = post.find('div', class_='price').text.strip() except: try: price = post.find('div', class_ = 'large-auto medium-auto cell').text.strip() except: price = "N/A" location = post.find('div', class_ = 'list-item-location').text.strip() desc = post.find('div', class_ = 'list-item-para').text.strip() try: tag = post.find('div', class_ = 'list-viewing-date').text.strip() except: tag = 'N/A' # 修复updated异常捕获的赋值错误 try: updated = post.find('div', class_ = 'list-update').text.strip() except: updated = 'N/A' try: client = post.find('p', class_ = 'client-name').text.strip() except: client = 'N/A' try: highlight = post.find('span', class_ = 'red-text').text.strip() except: highlight = 'N/A' # 参数字段不存在时统一赋值N/A years = sn = total_time = 'N/A' t = post.find_all('div',class_='list-other-dtl') for i in t: data = [tup.text.strip() for tup in i.find_all('li')] if len(data) >=3: years = data[0] sn = data[1] total_time = data[2] # 所有条目统一写入,不再放在参数遍历的循环内 temp.append([plane,price,location,years,sn,total_time,desc,tag,highlight,client,updated,link_full]) df=pd.DataFrame(temp,columns=["Plane","Price","location","Year","S/N","Total time","Description", "Tag", "Highlight", "Client", "Updated", "Link"]) # 编码选utf-8-sig避免打开乱码,同时去掉无用索引列 df.to_csv('/Users/xxx/avbuyer.csv', encoding='utf-8-sig', index=False)
调整说明
- 把数据写入逻辑从参数遍历块中移出,所有条目无论是否有参数都会被写入结果
- Price字段优先获取正常售价,拿不到时自动填充Wanted、Now Sold等状态值,无需额外单独存储状态列
- 修复了updated字段的赋值bug,增加参数长度校验,避免参数不足时报错
- 所有字段增加strip处理,清理多余的换行、空格
- 导出CSV时增加编码声明,避免乱码,同时去掉无用的索引列
内容的提问来源于stack exchange,提问作者Alistair
相关产品推荐
相关产品推荐

