You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫爬取sehat.com.pk商品仅能输出价格无法提取名称和链接如何解决

问题原因分析
  • 父容器选择错误:当前查找的col-md-12 pr-0 pl-0类不是单商品的包裹容器,遍历的对象不是商品节点,无法拿到对应属性
  • 链接提取逻辑错误:href是<a>标签的专属属性,对<div>标签取href只会返回空值或报错
  • 商品名提取逻辑错误:<img>是自闭合标签,本身没有内部文本,商品名称存储在<img>标签的alt属性中,不是text值
  • 请求缺少标识:目标站点有基础反爬校验,直接发送无UA的请求会返回异常页面,导致解析不到数据
  • 延时逻辑无效:requests.get()是同步阻塞方法,拿到响应后再执行sleep()完全不会影响请求结果,属于无效代码
修正后可运行代码
import requests
from bs4 import BeautifulSoup

url = 'https://sehat.com.pk/categories/Over-The-Counter-Drugs/Diarrhea-and-Vomiting-/'
# 增加请求头模拟浏览器访问,绕过基础反爬
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
r = requests.get(url, headers=headers)
soup = BeautifulSoup(r.content, 'html.parser')
# 选择正确的单商品父容器
content = soup.find_all('div', class_ = 'product_box')

base_url = "https://sehat.com.pk"
for item in content:
    # 从a标签取href,拼接完整路径
    link = base_url + item.find('a')['href']
    # 从img标签取alt属性作为商品名
    name = item.find('img', class_ = 'img-fluid')['alt'].strip()
    # 提取价格
    price = item.find('div', class_ = 'ProductPriceRating').text.strip()
    print(name, link, price)

内容的提问来源于stack exchange,提问作者user15290488

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 08:57:01