基于Python网页抓取实现两家线上超市商品价格对比的技术问询
Python线上超市商品价格对比函数优化方案
依赖准备
先安装必要的网页抓取库:
pip install requests beautifulsoup4
价格提取与对比实现
针对你提供的两个页面,先分析各自HTML结构定位价格元素,再实现健壮的提取与对比逻辑:
- Laughs超市页面:价格在带有
price类的span标签中 - Glomark超市页面:价格在带有
product-price类的div标签内
完整代码如下:
import requests from bs4 import BeautifulSoup def extract_price(url, site_type): try: # 模拟浏览器请求,避免被反爬拦截 response = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 按网站类型提取价格 if site_type == 'laughs': price_tag = soup.find('span', class_='price') if not price_tag: return None # 清理价格文本,提取纯数字 price_text = price_tag.get_text(strip=True).replace('Rs.', '').replace(',', '') elif site_type == 'glomark': price_tag = soup.find('div', class_='product-price') if not price_tag: return None price_text = price_tag.get_text(strip=True).replace('LKR', '').replace(',', '') else: return None return float(price_text) except Exception as e: print(f"提取价格出错:{str(e)}") return None def compare_prices(laughs_url, glomark_url): # 获取两家价格 laughs_price = extract_price(laughs_url, 'laughs') glomark_price = extract_price(glomark_url, 'glomark') # 检查数据完整性 if laughs_price is None or glomark_price is None: print("无法获取完整价格,对比终止") return # 输出对比结果 print(f"Laughs椰子价格:Rs.{laughs_price:.2f}") print(f"Glomark椰子价格:Rs.{glomark_price:.2f}") if laughs_price < glomark_price: print("Laughs的椰子更便宜") elif glomark_price < laughs_price: print("Glomark的椰子更便宜") else: print("两家椰子价格一致") # 执行对比 laughs_coconut = 'https://scrape-sm1.github.io/site1/COCONUT%20market1super.html' glomark_coconut = 'https://glomark.lk/coconut/p/11624' compare_prices(laughs_coconut, glomark_coconut)
核心优化说明
- 反爬规避:添加User-Agent头模拟正常浏览器请求
- 异常处理:捕获HTTP错误、元素查找失败等情况,避免程序崩溃
- 价格清洗:移除货币符号和千位分隔符,确保数值转换准确
- 模块化拆分:提取逻辑独立成函数,后续新增其他超市只需扩展
extract_price分支
内容的提问来源于stack exchange,提问作者lahiru pramuditha
相关产品推荐
相关产品推荐

