如何使用BeautifulSoup爬取多页数据 代码无输出问题排查求助
问题排查及修复方案
- 问题1:分页URL格式错误
你当前的请求URL拼接后是https://www.avbuyer.com/aircraft/private-jets=1,该路径不存在,服务端会返回404错误,无有效页面内容可解析。正确的分页参数应该以查询参数形式拼接,修改为https://www.avbuyer.com/aircraft/private-jets?page={page}即可。 - 问题2:缺少请求状态校验
当前代码没有判断请求是否成功,直接解析返回内容,一旦请求失败(404/500/反爬拦截等),后续解析自然没有结果。建议先打印请求状态码和返回内容,确认拿到了正确的页面源码再做解析。 - 问题3:元素选择器需匹配实际页面结构
如果修复前两个问题后依然无输出,需要确认你抓取的页面中,列表项的class是否真的是listing-item premium,部分站点会针对普通列表和付费推荐列表设置不同的class,你可以先打印soup.prettify()查看返回的源码结构,调整选择器即可。另外字段查找前建议加非空判断,避免属性不存在导致代码报错中断。
修正后可运行的参考代码
import requests from bs4 import BeautifulSoup # 补全User-Agent信息避免被反爬拦截 headers= {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'} # 要爬多页调整range结束值即可,比如range(1,10)就是爬1到9页 for page in range(1,2): # 修正URL格式,使用标准查询参数拼接分页 url = "https://www.avbuyer.com/aircraft/private-jets?page={page}".format(page=page) response = requests.get(url, headers=headers) # 先校验请求状态 if response.status_code != 200: print(f"第{page}页请求失败,状态码:{response.status_code}") continue soup = BeautifulSoup(response.content, 'html.parser') # 调整选择器匹配所有列表项,不需要限定premium类 postings = soup.find_all('div', class_ = 'listing-item') if not postings: print(f"第{page}页未找到匹配的列表项,请检查选择器或页面返回内容") continue for post in postings: # 字段查找前加非空判断,避免报错 link_ele = post.find('a', class_ = 'more-info') link_full = 'https://www.avbuyer.com' + link_ele.get('href') if link_ele else '' plane_ele = post.find('h2', class_ = 'item-title') plane = plane_ele.text.strip() if plane_ele else '' price_ele = post.find('div', class_ = 'price') price = price_ele.text.strip() if price_ele else '' location_ele = post.find('div', class_ = 'list-item-location') location = location_ele.text.strip() if location_ele else '' print(location, plane, price, link_full)
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

