You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何获取网页爬取的正确URL?以耐克巴西站点为例

爬取运动鞋价格:获取目标API接口的通用方法

问题背景

  • 已开发运动鞋价格爬取机器人,成功通过修正后的API接口获取Vans巴西站点商品价格:
    • Vans原商品URL:https://www.vans.com.br/tenis-ultrarange-rapidweld-black-white/p/1003500430051U?gad_source=1
    • 可用API接口:https://www.vans.com.br/arezzocoocc/v2/vans/products/1003500430051U/dynamic-product-fields?fields=DYNAMIC_FIELDS_PDP
  • 处理耐克巴西站点商品时,自行修改的接口无效:
    • 耐克原商品URL:https://www.nike.com.br/tenis-nike-pegasus-40-masculino-025803.html?cor=ID
    • 无效接口:https://www.nike.com.br/tenis-nike-pegasus-40-masculino-025803/dynamic-product-fields?fields=DYNAMIC_FIELDS_PDP

当前爬取代码

import requests
import smtplib
import email.message
import ssl

headers = {
    "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:121.0) Gecko/20100101 Firefox/121.0"
}
def get_product_data(url, number):
    try:
        response = requests.get(url, headers=headers)
        response.raise_for_status()
        data = response.json()
        return next((c for c in data.get("colorOptions", []) if c["code"] == number), None)
    except requests.RequestException as e:
        print(f"Error fetching product data for {number}: {e}")
        return None

def send_email(product_name, product_price, receiver_email):
    subject = "Price Drop Alert!"
    body_msg = f'''The price of {product_name} has dropped to {product_price}.'''
    message = f"Subject: {subject}\n\n{body_msg}"

    sender = '<REDACTED>'
    password = '<REDACTED>'
    receiver = '<REDACTED>'

    context = ssl.create_default_context()

    with smtplib.SMTP_SSL('smtp.gmail.com', 465, context=context) as server:
        server.login(sender, password)

        server.sendmail(sender, receiver, message.encode('utf-8'))

number1 = "1003500430051U"
number2 = "1002001070011U"

url1 = f"https://www.vans.com.br/arezzocoocc/v2/vans/products/{number1}/dynamic-product-fields?fields=DYNAMIC_FIELDS_PDP"
url2 = f"https://www.vans.com.br/arezzocoocc/v2/vans/products/{number2}/dynamic-product-fields?fields=DYNAMIC_FIELDS_PDP"

product1 = get_product_data(url1, number1)
product2 = get_product_data(url2, number2)

if product1 and "price" in product1:
    productprice1 = product1["price"]["value"]
    print(product1["name"], productprice1)

if product2 and "price" in product2:
    print(product2["name"], product2["price"]["value"])

try:
    data1 = requests.get(url1, headers=headers).json()
    data2 = requests.get(url2, headers=headers).json()
except requests.RequestException as e:
    print(f"Error fetching product data: {e}")

for c in data1["colorOptions"]:
    if c["code"] == number1:
        productprice1 = data1["price"]["value"]
        print(c["name"], productprice1)
        break

for c in data2["colorOptions"]:
    if c["code"] == number2:
        print(c["name"], data2["price"]["value"])
        break

if productprice1 and productprice1 < 600:
    send_email(product1["name"], productprice1, 'vini.damatta17@gmail.com')

通用获取爬取API接口的方法

1. 浏览器开发者工具抓包

  • 打开商品页面,按F12启动开发者工具,切换到Network标签
  • 刷新页面,筛选XHR/Fetch类型请求:这类请求通常是加载动态商品数据的API接口
  • 排查返回JSON格式的请求,查看Response内容是否包含价格、商品名称等所需字段,记录对应的接口URL、请求头和参数

2. 分析页面源码找线索

  • 在Elements标签中搜索关键词(如price、api、productId)
  • 部分站点会在页面<script>标签中嵌入API模板或商品唯一标识,比如Vans接口中的商品ID就来自原URL的路径参数
  • 检查是否存在window.productInfo这类全局变量,其中可能直接包含API接口地址或商品数据

3. 模拟浏览器请求参数

  • 多数站点会校验请求头,确保requests请求携带完整的User-Agent、Referer(原商品页面URL)等字段
  • 若接口需要Cookie或Authorization令牌,直接从抓包结果中复制对应参数到请求头

4. Nike巴西站点临时适配方案

针对你提供的Nike商品URL:

  • 抓包后可发现其商品数据接口通常为https://www.nike.com.br/api/product/{product_id},其中product_id可从原URL的数字部分提取(如示例中的025803)
  • 验证接口返回的JSON数据,提取price字段对应的价格值

5. 爬取注意事项

  • 设置请求间隔(如time.sleep(2)),避免频繁请求导致IP被封禁
  • 查看站点robots.txt规则,确认是否允许爬取目标数据
  • 若遇到验证码、动态令牌等反爬机制,可使用Selenium或Playwright模拟完整浏览器行为

内容的提问来源于stack exchange,提问作者Leandro Vinícius da Mata Silva

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:55:15