You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行Etsy分析Jupyter Notebook时Pandas apply()函数报错求助

Troubleshooting Your Etsy Data Acquisition Notebook Error

Hey there! Let's dig into why that apply() line is throwing errors in your notebook. First, let's recap what you're working with:

You're running a notebook focused on Etsy data acquisition, and this line is failing:

# Apply function and store multiple page urls in column called page_nums
catdf['page_nums'] = catdf.apply(get_page2, axis=1)

From the error screenshots you shared, I can spot common pitfalls that cause issues with df.apply() when using axis=1. Let's break down the likely causes and fixes:

1. Your get_page2 function isn't handling row inputs correctly

When you use apply(axis=1), pandas passes each row as a Series object to your function. If get_page2 isn't set up to accept this row as a parameter, it'll throw an error right away.

Fix: Double-check your function definition to make sure it accepts a row argument. For example:

def get_page2(row):
    # Grab the base category URL from the row
    base_url = row["category_url"]
    # Add your logic here to calculate total pages and generate page URLs
    total_pages = calculate_total_pages(base_url)  # Replace with your actual logic
    return [f"{base_url}?page={page_num}" for page_num in range(1, total_pages + 1)]

2. Missing or invalid data in your catdf DataFrame

If some rows in catdf are missing the fields get_page2 depends on (like a category URL), the function will crash when trying to access that data.

Fix: First, check for missing values and clean your data:

# Check which columns have missing values
print(catdf.isnull().sum())

# Drop rows that don't have the critical category URL field
catdf = catdf.dropna(subset=["category_url"])

3. Network/Scraping Issues (if get_page2 fetches Etsy pages)

If get_page2 is scraping Etsy to get total pages, you might be hitting anti-scraping measures, timeouts, or changes to Etsy's page structure.

Fix: Add error handling, retries, and respect Etsy's rules:

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def get_page2(row):
    # Set up a session with retries to handle temporary network issues
    session = requests.Session()
    retry_strategy = Retry(
        total=3,
        backoff_factor=0.5,
        status_forcelist=[429, 500, 502, 503, 504]
    )
    adapter = HTTPAdapter(max_retries=retry_strategy)
    session.mount("https://", adapter)
    
    try:
        # Fetch the page with a timeout to avoid hanging
        response = session.get(row["category_url"], timeout=10)
        response.raise_for_status()  # Raise an error for HTTP status codes like 404/500
        
        # Parse the page to get total pages (adjust selector to match Etsy's current structure)
        soup = BeautifulSoup(response.text, "html.parser")
        total_pages_text = soup.find("span", class_="pagination-count").text
        total_pages = int(total_pages_text.split()[-1])
        
        # Generate all page URLs
        return [f"{row['category_url']}?page={num}" for num in range(1, total_pages + 1)]
    
    except Exception as e:
        # Print a helpful error message instead of crashing the whole apply
        print(f"Failed to process row {row.name}: {str(e)}")
        return None  # Or return an empty list if you prefer

Note: Etsy has strict anti-scraping policies. Make sure you're following their robots.txt and not making too many requests too quickly—add delays if needed!

4. Invalid return type from get_page2

Pandas needs the function to return a value that can be stored in a DataFrame column (like a list, string, or number). If your function returns a complex object or nothing at all, you'll get an error.

Fix: Ensure get_page2 always returns a consistent type (e.g., a list of URLs, or None/empty list on failure).

Quick Debugging Tip

Before running the full apply(), test get_page2 on a single row to isolate the issue:

# Test with the first row in your DataFrame
test_row = catdf.iloc[0]
print(get_page2(test_row))

This will show you exactly where the function is failing without running it on the entire dataset.

内容的提问来源于stack exchange,提问作者user8022517

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:06:20