You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup无法提取商品href属性问题咨询

排查结果&修复方案
  • 第一个问题:requests.get传参错误
    你写的r2 = requests.get(url2, headers)是错误用法,headers属于关键字参数,必须指定参数名才能生效,你这么写相当于把headers字典传给了第二个位置参数params,请求头里的UA还是requests默认值,直接触发网站反爬拦截,返回的不是正常商品页面,自然解析不到对应元素。
  • 第二个问题:代码存在语法&逻辑错误
    1. 没有提前初始化link_list列表,直接调用append方法会触发NameError报错
    2. 循环内缩进错误,link_list.append那行缩进不符合Python规范,会直接报语法错误无法运行
    3. 最后一行单独写的link_list没有实际作用
  • 第三个问题:缺少基础反爬规避措施
    该网站除了UA校验外,还会校验Accept、Accept-Language等常规请求头,缺少这些头即使UA正确也可能被拦截。

修复后可运行代码

import requests
from bs4 import BeautifulSoup as bs4

# 初始化存储列表
link_list = []
url2 = "https://www.vidri.com.sv/catalogo/07040101/Taladros-y-atornilladores-inalambricos.html"

# 完善请求头,符合浏览器正常请求规则
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3",
    "Referer": "https://www.vidri.com.sv/"
}

# 关键字参数传递headers,加超时设置避免长时间卡住
r2 = requests.get(url2, headers=headers, timeout=10)
# 校验请求是否成功,状态码非200会直接抛出异常
r2.raise_for_status()

soup2 = bs4(r2.content, "lxml")
products_container = soup2.find("div", class_= "products__container")

if products_container:
    for item in products_container.find_all("div", class_= "productCard"):
        for sku in item.find_all("a", href = True):
            link_list.append(sku["href"])
    print("提取到的商品链接:", link_list)
else:
    print("未找到商品容器,页面返回异常,可打印r2.text查看实际返回内容确认反爬类型")

如果运行后还是提示未找到商品容器,说明触发了更严格的反爬机制(比如JS动态渲染、Cookie校验、人机验证),可以改用Selenium、Playwright等无头浏览器工具模拟真实用户访问。

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 01:45:04