You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python API爬虫多格式价格提取异常:无法获取区间价格、输出重复问题求助

Fixing Price Extraction & Duplicate Output Issues in Your API Scraper

Let's break down and fix the two issues you're facing with your price extraction code:


Issue 1: Price is None for APIs returning price_min/price_max

Your current logic has two critical flaws here:

  1. You’re using zip(WANTED_PRICES, WANTED_avgPRICES) to loop through pairs of fields, which means you’re checking one direct average field and one min/max field at a time instead of first checking all direct average fields, then verifying if both min/max prices exist together.
  2. The elif result_mm in result check only verifies one of the min/max fields exists, but you need both to calculate a valid average. If your loop hits a single min/max field without the other, price stays None.

Issue 2: Duplicate Outputs

The zip loop runs for the length of the shorter list (2, since WANTED_avgPRICES has 2 elements). This means each result gets processed twice, leading to duplicate print statements for every data entry.


Fixed Code

Here's the revised version of your parse_slugs method with fixes for both issues:

from decimal import Decimal
import json

def parse_slugs(self, response, slug_id, crawled_url):
    summary = json.loads(response.body)
    results = summary.get('results', [])
    for result in results:
        price = None
        # First check for direct average price fields
        for avg_field in WANTED_PRICES:
            if avg_field in result:
                price = Decimal(result[avg_field])
                break  # Stop checking once we find a valid field
        
        # If no direct average field found, check for BOTH min/max prices
        if price is None:
            if 'price_min' in result and 'price_max' in result:
                minp = result['price_min']
                maxp = result['price_max']
                # Fix: Convert each value to Decimal BEFORE adding (avoids string concatenation)
                price = (Decimal(minp) + Decimal(maxp)) / 2
        
        # Handle date field more safely
        date = None
        if 'published_date' in result:
            date = result['published_date'].split(' ')[0]
        elif 'published_Date' in result:
            date = result['published_Date'].split(' ')[0]
        # Optional: Add a fallback if neither date field exists
        # else:
        #     date = "Unknown"
        
        # Print only if we have valid date/price data
        if date:
            print(slug_id, date, price if price is not None else "No valid price found")

Key Fixes Explained

  1. Direct Price Field Priority: We loop through WANTED_PRICES first and break immediately when we find a matching field, ensuring we don’t waste time checking unnecessary fields.
  2. Valid Min/Max Check: We only calculate the average if both price_min and price_max exist in the result, preventing invalid calculations and None values.
  3. Decimal Calculation Fix: Your original code did Decimal(minp + maxp) which would concatenate strings (common in API responses) before converting to Decimal. We now convert each value individually first for accurate numerical math.
  4. Eliminated Duplicates: By removing the zip loop and processing each result exactly once, we avoid duplicate print outputs entirely.
  5. Safer Date Handling: We added explicit elif for the camelCase date field and optional fallback logic to avoid unexpected KeyErrors if neither date field exists.

内容的提问来源于stack exchange,提问作者Shruthi Ravishankar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 11:57:33