You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫数据转换报错修复:非英文数字与空值处理

Fixing Scrapy Parsing Errors: Handling Non-Numeric Data and None Values

Hey there! Let's work through those two frustrating parsing issues you're hitting in your Scrapy crawler. These are super common when dealing with real-world e-commerce HTML—data is rarely clean, so we need to build more robust parsing logic to handle edge cases.

First Error: ValueError from Invalid Numeric Conversion

Your initial crash came from trying to convert a messy string directly to an integer. The string had newlines, spaces, and a non-English percent symbol (٪), which int() couldn't parse:

ValueError: invalid literal for int() with base 10: '\n ٪۲۹\n '

Cleaning the string was the right first step, but we needed to add a safety net for when elements don't exist on the page.

Second Error: AttributeError from NoneType.strip()

When you added strip(), you hit a new issue: sometimes your XPath queries return None (like when a product doesn't display a discount or original price). Calling strip() on None immediately throws an error:

AttributeError: 'NoneType' object has no attribute 'strip'

Robust Solution: Handle None and Clean Data Safely

Here's how to rewrite your parsing code to fix both problems. We'll use defensive checks and a reusable helper function to make the code more maintainable:

# Helper function to safely clean and convert messy numeric strings
def safe_numeric_conversion(raw_value, replace_map=None):
    # Return None if input is empty or None to avoid method calls on invalid types
    if not raw_value:
        return None
    
    # Step 1: Remove extra whitespace (newlines, spaces, tabs)
    cleaned_str = raw_value.strip()
    
    # Step 2: Replace unwanted characters (like Persian % or commas)
    if replace_map:
        for char, replacement in replace_map.items():
            cleaned_str = cleaned_str.replace(char, replacement)
    
    # Step 3: Try converting to int first, then float, return None if both fail
    try:
        return int(cleaned_str)
    except ValueError:
        try:
            return float(cleaned_str)
        except ValueError:
            return None

# Parse each product field with safety checks
# Discount percentage (handle Persian % symbol)
discount_raw = product.xpath("//div[@class='c-price__discount-oval']/span/text()").get(default='')
discounted_percent = safe_numeric_conversion(discount_raw, {'٪': ''})

# Discounted price (handle commas as thousand separators)
discounted_price_raw = product.xpath("//div[@class='c-price__value-wrapper']/text()").get(default='')
discounted_price = safe_numeric_conversion(discounted_price_raw, {',': ''})

# Original price (handle commas)
original_price_raw = product.xpath("//div[@class='c-price__value c-price__value--plp']/del/text()").get(default='')
original_price = safe_numeric_conversion(original_price_raw, {',': ''})

# Apply your filters only to valid data
if discounted_percent is not None and discounted_percent >= 10:
    # Don't forget to apply the same safe logic to your rating field!
    # Example: rating = safe_numeric_conversion(rating_raw, {'/5': ''})
    # if rating is not None and rating >= 3.5:
    #     process and store the product data

Key Improvements:

  • get(default=''): Ensures XPath queries never return None—we get an empty string instead if the element is missing, preventing AttributeError.
  • Reusable Helper Function: Centralizes cleaning and conversion logic, so you don't repeat code for each field.
  • Defensive Checks: We only process products where the discount (and rating, once you add it) is valid and meets your thresholds.
  • Flexible Character Replacement: Use a dictionary to target specific unwanted characters, making it easy to adjust for other edge cases.

内容的提问来源于stack exchange,提问作者Shima Masaeli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:37:49