Python使用BeautifulSoup提取class内列表的网页爬虫问题求助
目标站点列表数据提取解决方案
你通过soup.select(".stockItemInfo")获取到的是所有匹配class的元素集合,每个元素对应一条独立的车辆信息区块,按以下步骤完成内容提取即可:
- 先修正代码中
headers内User-Agent的换行错误,避免请求发送失败 - 遍历所有stockItemInfo元素,按需提取内部的文本、标签属性等内容
- 将提取后的结构化数据汇总存储,支持导出为csv等格式
完整可运行代码
import bs4, requests import pandas as pd wagon_stock_url = 'https://parramattamg.com.au/up4053-961230-mg-hs-2020.html' # 修复User-Agent换行问题 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/96.0.4664.45 Safari/537.36' } response = requests.get(wagon_stock_url, headers=headers) soup = bs4.BeautifulSoup(response.text, 'html.parser') # 获取所有车辆信息区块 stock_item_list = soup.select(".stockItemInfo") extracted_data = [] for item in stock_item_list: current_item = {} # 提取车辆标题 title = item.find('h2') current_item['车辆名称'] = title.get_text(strip=True) if title else '无' # 提取价格 price = item.find(class_='price') current_item['售价'] = price.get_text(strip=True) if price else '无' # 提取所有配置列表项内容 specs = [li.get_text(strip=True) for li in item.find_all('li')] current_item['配置列表'] = specs # 如需提取整个区块的纯文本,可使用下行代码 # current_item['全部内容'] = item.get_text(strip=True, separator=' ') extracted_data.append(current_item) # 转为表格结构,可直接导出为csv df = pd.DataFrame(extracted_data) print(df) # 导出执行:df.to_csv('mg车辆库存数据.csv', index=False, encoding='utf-8-sig')
自定义提取说明
如果需要提取更细分的字段,可以用浏览器F12开发者工具定位目标内容的标签、class属性,调整find/find_all/select的参数即可,层级选择写法示例:
- 选中当前区块下class为
detail-row内的所有p标签:item.select('.detail-row p') - 选中当前区块下属性为
data-type="spec"的元素:item.find(attrs={'data-type':'spec'})
内容的提问来源于stack exchange,提问作者Kushal_70
相关产品推荐
相关产品推荐

