Scrapy爬虫数据转换报错修复:非英文数字与空值处理
Hey there! Let's work through those two frustrating parsing issues you're hitting in your Scrapy crawler. These are super common when dealing with real-world e-commerce HTML—data is rarely clean, so we need to build more robust parsing logic to handle edge cases.
First Error: ValueError from Invalid Numeric Conversion
Your initial crash came from trying to convert a messy string directly to an integer. The string had newlines, spaces, and a non-English percent symbol (٪), which int() couldn't parse:
ValueError: invalid literal for int() with base 10: '\n ٪۲۹\n '
Cleaning the string was the right first step, but we needed to add a safety net for when elements don't exist on the page.
Second Error: AttributeError from NoneType.strip()
When you added strip(), you hit a new issue: sometimes your XPath queries return None (like when a product doesn't display a discount or original price). Calling strip() on None immediately throws an error:
AttributeError: 'NoneType' object has no attribute 'strip'
Robust Solution: Handle None and Clean Data Safely
Here's how to rewrite your parsing code to fix both problems. We'll use defensive checks and a reusable helper function to make the code more maintainable:
# Helper function to safely clean and convert messy numeric strings def safe_numeric_conversion(raw_value, replace_map=None): # Return None if input is empty or None to avoid method calls on invalid types if not raw_value: return None # Step 1: Remove extra whitespace (newlines, spaces, tabs) cleaned_str = raw_value.strip() # Step 2: Replace unwanted characters (like Persian % or commas) if replace_map: for char, replacement in replace_map.items(): cleaned_str = cleaned_str.replace(char, replacement) # Step 3: Try converting to int first, then float, return None if both fail try: return int(cleaned_str) except ValueError: try: return float(cleaned_str) except ValueError: return None # Parse each product field with safety checks # Discount percentage (handle Persian % symbol) discount_raw = product.xpath("//div[@class='c-price__discount-oval']/span/text()").get(default='') discounted_percent = safe_numeric_conversion(discount_raw, {'٪': ''}) # Discounted price (handle commas as thousand separators) discounted_price_raw = product.xpath("//div[@class='c-price__value-wrapper']/text()").get(default='') discounted_price = safe_numeric_conversion(discounted_price_raw, {',': ''}) # Original price (handle commas) original_price_raw = product.xpath("//div[@class='c-price__value c-price__value--plp']/del/text()").get(default='') original_price = safe_numeric_conversion(original_price_raw, {',': ''}) # Apply your filters only to valid data if discounted_percent is not None and discounted_percent >= 10: # Don't forget to apply the same safe logic to your rating field! # Example: rating = safe_numeric_conversion(rating_raw, {'/5': ''}) # if rating is not None and rating >= 3.5: # process and store the product data
Key Improvements:
get(default=''): Ensures XPath queries never returnNone—we get an empty string instead if the element is missing, preventingAttributeError.- Reusable Helper Function: Centralizes cleaning and conversion logic, so you don't repeat code for each field.
- Defensive Checks: We only process products where the discount (and rating, once you add it) is valid and meets your thresholds.
- Flexible Character Replacement: Use a dictionary to target specific unwanted characters, making it easy to adjust for other edge cases.
内容的提问来源于stack exchange,提问作者Shima Masaeli

