Python亚马逊爬虫提取子类目时遇TypeError问题求助
First up, let's tackle the TypeError: 'NoneType' object is not callable error stopping your crawl—it's the immediate roadblock, so we'll fix that first.
1. Fixing the TypeError
The error pops up because page is None when you try to call page.find_all(). This means your make_request(line) function is returning a None value for page, which usually happens for one of these reasons:
- The HTTP request to Amazon failed (e.g., you got a 403 Forbidden response due to anti-scraping measures, or a network error).
- The response HTML was empty or invalid, so BeautifulSoup couldn't parse it into a usable object.
Quick Fixes:
- Add a validation check right after calling
make_requestto skip broken pages:page, html = make_request(line) if not page: log(f"Failed to load/parse page: {line}") continue # Move to the next URL in your list - Audit your
make_requestfunction: Ensure it sets proper request headers (like a realUser-Agentto mimic a browser) and only returns a BeautifulSoup object if the response status code is 200. Amazon aggressively blocks scrapers without legitimate headers.
2. Alternative Methods to Extract Amazon Subcategories
Your current selectors (mm-column, mm-category-list) are likely outdated—Amazon frequently updates its page structure, so those classes may no longer exist on the target workout clothes page. Here are more reliable approaches:
Approach 1: Target Links with Category Node IDs
Amazon subcategory URLs almost always include a node= parameter (e.g., /node/11444071011). You can filter directly for these links:
# Find all links containing a category node reference subcategory_links = page.find_all( "a", href=lambda href: href and ("node=" in href or "/node/" in href) ) # Deduplicate links (Amazon often repeats category links across the page) unique_links = list({link["href"] for link in subcategory_links}) count = 0 for link in unique_links: enqueue_url(link) count += 1
Approach 2: Target the Left Navigation Bar
Most Amazon category pages have a left sidebar dedicated to subcategories. You can target this container directly (use your browser's dev tools to confirm current IDs/classes):
# Locate the left navigation sidebar (current ID as of 2024 is "leftNav") left_nav = page.find("div", id="leftNav") if left_nav: # Find all indented list items (common for nested subcategories) subcategory_items = left_nav.find_all( "li", class_=lambda c: c and "s-ref-indent" in c ) count = 0 for item in subcategory_items: link = item.find("a") if link and "href" in link.attrs: enqueue_url(link["href"]) count += 1
Pro Tip: Use Browser Dev Tools
Always inspect the target page with your browser's "Inspect Element" tool to confirm current class names/IDs. For your workout clothes page, you'll notice subcategories are nested in lists with classes like a-unordered-list a-nostyle a-vertical s-ref-indent-one—adjust your selectors to match what's actually present on the page.
内容的提问来源于stack exchange,提问作者ryy77

