BeautifulSoup爬取avbuyer私人飞机数据导出CSV时字段提取求助
问题原因
你代码的失效是几个语法和逻辑错误共同导致的,具体如下:
- 提取
year/sn/time时,你调用了.text取文本但没有赋值回变量,最终存入DataFrame的是BeautifulSoup的DOM元素对象而非文本值 sn和time提取时混淆了find和find_all:find返回单个元素不支持下标索引,只有find_all返回的列表才能用[2]取第三个ul节点- 提取描述字段时类名拼写错误,
classs_多写了一个s,导致无法匹配到对应节点 - 无限制的
try/except会吞掉所有报错,你无法定位具体是哪一步出了问题
修正后的可运行代码
import requests from bs4 import BeautifulSoup import pandas as pd url = 'https://www.avbuyer.com/aircraft/private-jets' headers = { "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } # 初始化空列表存数据,比每次append df性能高很多 data_list = [] while True: page = requests.get(url, headers=headers) page.raise_for_status() # 接口报错直接抛出,方便排查 soup = BeautifulSoup(page.text, 'lxml') postings = soup.find_all('div', class_ = 'listing-item premium') for post in postings: try: link = post.find('a', class_ = 'more-info').get('href') link_full = 'https://www.avbuyer.com'+ link plane = post.find('h2', class_ = 'item-title').text.strip() price = post.find('div', class_ = 'price').text.strip() location = post.find('div', class_ = 'list-item-location').text.strip() # 修正find_all + 赋值text info_ul = post.find_all('ul', class_ = 'fa-no-bullet clearfix')[2] year = info_ul.find_all('li')[0].text.strip() sn = info_ul.find_all('li')[1].text.strip() time = info_ul.find_all('li')[2].text.strip() # 修正class拼写错误 desc = post.find('div', class_ = 'list-item-para').text.strip() tag = post.find('div', class_ = 'list-viewing-date').text.strip() updated = post.find('div', class_ = 'list-update').text.strip() # 加入列表 data_list.append({ 'Plane':plane, 'Year':year, 'S/N':sn, 'Total Time':time, 'Price':price, 'Location':location, 'Description':desc, 'Tag':tag, 'Last updated':updated, 'Link':link_full }) except Exception as e: # 打印错误方便排查,不需要可以删掉该行 continue # 下一页判断,避免最后一页报错 next_page_tag = soup.find('a', {'rel':'next'}) if not next_page_tag: break next_page = next_page_tag.get('href') url = 'https://www.avbuyer.com' + next_page # 列表一次性转df,性能更高 df = pd.DataFrame(data_list) df.to_csv('/Users/xxx/avbuyer.csv', index=False, encoding='utf-8-sig')
额外优化点
- 加了
User-Agent请求头,避免被网站反爬拦截 - 改用列表存数据最后一次性转DataFrame,比循环append df性能高10倍以上
- 增加了下一页存在性判断,避免最后一页没有下一页按钮时抛出报错
- 增加
.strip()处理文本首尾的空白字符、换行符 - 导出CSV指定
utf-8-sig编码,避免打开时乱码
内容的提问来源于stack exchange,提问作者Alistair
相关产品推荐
相关产品推荐

