You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Shopify API分页获取产品数据中途停止问题排查

Shopify产品API分页获取中断问题的排查与修复

可能的中断原因

  • 未处理API限流:Shopify私有APP的API调用限制为60秒内最多40次请求,触发限流时会返回429状态码,原代码直接终止循环,导致获取流程中断。
  • 高频请求触发限制:连续无间隔的请求极易触发限流机制,没有在请求间加入合理延迟。
  • 缺失变体容错处理:若某产品无变体,原代码中product['variants'][0]会抛出索引错误,直接终止程序。

修复后的代码

import requests
import csv
import time
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

API_KEY = "key"
PASSWORD = "password"
SHOP_HANDLE = "name"
API_VERSION = '2023-07'
CSV_FILE_PATH = 'product_inventory_item_ids.csv'

def create_session():
    # 创建带重试机制的会话,自动处理限流和服务器错误
    session = requests.Session()
    retry_strategy = Retry(
        total=5,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET"]
    )
    adapter = HTTPAdapter(max_retries=retry_strategy)
    session.mount("https://", adapter)
    session.auth = (API_KEY, PASSWORD)
    return session

def get_all_products():
    page_info = None
    all_products = []
    session = create_session()

    while True:
        url = f"https://{SHOP_HANDLE}.myshopify.com/admin/api/{API_VERSION}/products.json?limit=250"
        if page_info:
            url += f"&page_info={page_info}"

        try:
            response = session.get(url)
            response.raise_for_status()  # 主动抛出HTTP错误
        except requests.exceptions.RequestException as e:
            print(f'请求出错: {str(e)}')
            # 针对429限流,按返回的Retry-After头指定时间等待
            if hasattr(e, 'response') and e.response.status_code == 429:
                retry_after = int(e.response.headers.get('Retry-After', 10))
                print(f'触发限流,等待{retry_after}秒后重试...')
                time.sleep(retry_after)
                continue
            break

        products_data = response.json()
        products = products_data.get('products', [])
        
        if not products:
            break

        all_products.extend(products)
        page_info = products_data.get('next_page_info')

        if not page_info:
            break

        # 请求间加入小延迟,降低限流风险
        time.sleep(0.5)

    return all_products

def main():
    with open(CSV_FILE_PATH, 'w', newline='', encoding='utf-8') as csv_file:
        csv_writer = csv.writer(csv_file)
        csv_writer.writerow(['Product ID', 'Inventory Item ID'])
        
        for product in get_all_products():
            product_id = product['id']
            # 检查产品是否存在变体,避免索引错误
            if product.get('variants') and len(product['variants']) > 0:
                inventory_item_id = product['variants'][0]['inventory_item_id']
                csv_writer.writerow([product_id, inventory_item_id])
            else:
                print(f'产品ID {product_id} 无变体,已跳过')

    print('Proceso completado. Los datos se han guardado en el archivo CSV.')

if __name__ == "__main__":
    main()

关键优化说明

  • 限流自动重试:通过urllib3.Retry实现自动重试429、5xx类错误,配合Retry-After头信息精准等待,避免手动处理的繁琐。
  • 请求频率控制:每次请求后加入0.5秒延迟,降低触发限流的概率。
  • 异常容错增强:捕获所有请求异常,针对限流做特殊处理,避免直接终止循环。
  • 变体存在性检查:提前判断产品是否有变体,防止因索引错误导致程序崩溃。

内容的提问来源于stack exchange,提问作者Sofia Perez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 18:34:57