运行Etsy分析Jupyter Notebook时Pandas apply()函数报错求助
Hey there! Let's dig into why that apply() line is throwing errors in your notebook. First, let's recap what you're working with:
You're running a notebook focused on Etsy data acquisition, and this line is failing:
# Apply function and store multiple page urls in column called page_nums catdf['page_nums'] = catdf.apply(get_page2, axis=1)
From the error screenshots you shared, I can spot common pitfalls that cause issues with df.apply() when using axis=1. Let's break down the likely causes and fixes:
1. Your get_page2 function isn't handling row inputs correctly
When you use apply(axis=1), pandas passes each row as a Series object to your function. If get_page2 isn't set up to accept this row as a parameter, it'll throw an error right away.
Fix: Double-check your function definition to make sure it accepts a row argument. For example:
def get_page2(row): # Grab the base category URL from the row base_url = row["category_url"] # Add your logic here to calculate total pages and generate page URLs total_pages = calculate_total_pages(base_url) # Replace with your actual logic return [f"{base_url}?page={page_num}" for page_num in range(1, total_pages + 1)]
2. Missing or invalid data in your catdf DataFrame
If some rows in catdf are missing the fields get_page2 depends on (like a category URL), the function will crash when trying to access that data.
Fix: First, check for missing values and clean your data:
# Check which columns have missing values print(catdf.isnull().sum()) # Drop rows that don't have the critical category URL field catdf = catdf.dropna(subset=["category_url"])
3. Network/Scraping Issues (if get_page2 fetches Etsy pages)
If get_page2 is scraping Etsy to get total pages, you might be hitting anti-scraping measures, timeouts, or changes to Etsy's page structure.
Fix: Add error handling, retries, and respect Etsy's rules:
import requests from bs4 import BeautifulSoup from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry def get_page2(row): # Set up a session with retries to handle temporary network issues session = requests.Session() retry_strategy = Retry( total=3, backoff_factor=0.5, status_forcelist=[429, 500, 502, 503, 504] ) adapter = HTTPAdapter(max_retries=retry_strategy) session.mount("https://", adapter) try: # Fetch the page with a timeout to avoid hanging response = session.get(row["category_url"], timeout=10) response.raise_for_status() # Raise an error for HTTP status codes like 404/500 # Parse the page to get total pages (adjust selector to match Etsy's current structure) soup = BeautifulSoup(response.text, "html.parser") total_pages_text = soup.find("span", class_="pagination-count").text total_pages = int(total_pages_text.split()[-1]) # Generate all page URLs return [f"{row['category_url']}?page={num}" for num in range(1, total_pages + 1)] except Exception as e: # Print a helpful error message instead of crashing the whole apply print(f"Failed to process row {row.name}: {str(e)}") return None # Or return an empty list if you prefer
Note: Etsy has strict anti-scraping policies. Make sure you're following their robots.txt and not making too many requests too quickly—add delays if needed!
4. Invalid return type from get_page2
Pandas needs the function to return a value that can be stored in a DataFrame column (like a list, string, or number). If your function returns a complex object or nothing at all, you'll get an error.
Fix: Ensure get_page2 always returns a consistent type (e.g., a list of URLs, or None/empty list on failure).
Quick Debugging Tip
Before running the full apply(), test get_page2 on a single row to isolate the issue:
# Test with the first row in your DataFrame test_row = catdf.iloc[0] print(get_page2(test_row))
This will show you exactly where the function is failing without running it on the entire dataset.
内容的提问来源于stack exchange,提问作者user8022517

