Web Scraping及BeautifulSoup下一页解析与航空数据CSV导出问题
问题原因
现有代码仅完成了第一页的数据提取、第二页的页面请求,没有对第二页及后续页面的列表数据做提取逻辑,也没有循环遍历所有分页,最终CSV仅存储第一页数据。另外存在两个可优化点:
- 定义了
headers请求头但未传入requests.get方法,容易触发站点反爬策略 - 两次页面解析使用了不同的解析器(
html.parser和lxml),可能出现解析结果不一致的问题
修复后完整代码
import requests from bs4 import BeautifulSoup import pandas as pd import time headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'} base_url = 'https://www.avbuyer.com' current_url = 'https://www.avbuyer.com/aircraft/private-jets' temp = [] while True: # 请求当前页 response = requests.get(current_url, headers=headers) soup = BeautifulSoup(response.content, 'lxml') # 提取当前页所有列表数据 postings = soup.find_all('div', class_ = 'listing-item premium') for post in postings: link = post.find('a', class_ = 'more-info').get('href') link_full = base_url + link plane = post.find('h2', class_ = 'item-title').text.strip() price = post.find('div', class_ = 'price').text.strip() location = post.find('div', class_ = 'list-item-location').text.strip() desc = post.find('div', class_ = 'list-item-para').text.strip() try: tag = post.find('div', class_ = 'list-viewing-date').text.strip() except: tag = 'N/A' updated = post.find('div', class_ = 'list-update').text.strip() t = post.find_all('div',class_='list-other-dtl') for i in t: data = [tup.text.strip() for tup in i.find_all('li')] years = data[0] s = data[1] total_time = data[2] temp.append([plane,price,location,years,s,total_time,desc,tag,updated,link_full]) # 查找下一页链接,没有则退出循环 next_tag = soup.find('a', {'rel':'next'}) if not next_tag: break next_page = next_tag.get('href') current_url = base_url + next_page # 加延时避免被封 time.sleep(2) # 导出CSV df = pd.DataFrame(temp, columns=["plane","price","location","Year","S/N","Totaltime","Description","Tag","Last Updated","link"]) df.to_csv('/Users/xxx/avbuyer.csv', index=False, encoding='utf-8-sig')
注意事项
- 代码中加入了
time.sleep(2)的请求间隔,避免请求频率过高被站点封禁IP,可根据实际情况调整间隔时长 - 导出CSV时指定了
encoding='utf-8-sig',避免Windows系统下打开CSV出现乱码 - 如果需要爬取非置顶的普通私人机信息,可以将列表选择器从
listing-item premium调整为匹配所有listing-item类的div - 运行环境不会影响执行结果,Python 3.8.12版本完全兼容所有依赖库,Chrome浏览器未参与爬虫请求流程,和运行结果无关
内容的提问来源于stack exchange,提问作者Alistair
相关产品推荐
相关产品推荐

