Python嵌套标签爬取问题:二手车网站价格与车型提取
解决easycar.tw二手车价格与车型提取问题
原代码问题分析
你编写的get_basic_info函数存在两处问题:
- 语法错误:
append语句缺少闭合括号,运行会直接报错 - 逻辑错误:
find_all的参数传入逻辑混乱,不需要嵌套调用item.find_all
修正后的提取方案
结合提供的HTML结构,我们可以直接定位span.price提取价格,再通过h5标签的剩余文本提取车型名称。
完整代码示例
from requests import get from bs4 import BeautifulSoup # 补充请求头避免被反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } def get_basic_info(content_list): basic_info = [] for item in content_list: # 提取价格:定位span.price标签并获取文本 price_tag = item.find("span", {"class": "price"}) price = price_tag.get_text(strip=True) if price_tag else "无价格" # 提取车型:移除价格文本后,获取剩余的有效内容 model = item.get_text(strip=True).replace(price, "").strip() basic_info.append({ "价格": price, "车型": model }) return basic_info # 爬取1-17页(range左闭右开,18为终止值实际到17页) for page in range(1, 18): base_url = f"https://www.easycar.tw/carList.php?Action=search&show=col&lifting=desc&year=&year1=&page={page}" response = get(base_url, headers=headers) response.encoding = "utf-8" # 强制指定编码,避免中文乱码 html_soup = BeautifulSoup(response.text, 'html.parser') content_list = html_soup.find_all(attrs={"class": "caption"}) # 提取并打印单页车辆信息 car_info = get_basic_info(content_list) for info in car_info: print(f"价格: {info['价格']}") print(f"车型: {info['车型']}") print("---")
代码说明
- 价格提取:用
find精准定位单个价格标签,get_text(strip=True)自动去除文本首尾的空白字符 - 车型提取:先获取
h5标签的全部文本,再通过替换价格内容,得到纯车型名称 - 优化了URL拼接方式,使用f-string更简洁易读;添加编码设置解决中文乱码问题
输出示例
价格: 59.8 萬 车型: TOYOTA ALTIS --- 价格: 32.8 萬 车型: 2011 SUBARU FORESTER --- 价格: 108.8 萬 车型: 2017 LEXUS ES ---
内容的提问来源于stack exchange,提问作者Henry Liu
相关产品推荐
相关产品推荐

