开发维基分类爬虫时遍历列表触发TypeError,寻求技术解决方案
Fixing the TypeError in Your Wikipedia Category Scraper
Hey there! Let's break down why you're hitting that TypeError and get your scraper working properly.
What's Causing the Error?
The main issue is that not every <a> tag on the page has an href attribute. When you call item.get('href') on those tags, it returns None instead of a string. Later, when you check 'Category' in item, you're trying to run a string check on a None value—which triggers the TypeError.
Step-by-Step Fix & Improved Code
Here's a revised version of your code that fixes the error, handles relative URLs, and properly grabs the top 10 links from each category page:
import urllib.request from bs4 import BeautifulSoup # Use `with` statement to auto-close the file (safer than manual open/close) with open('url_list', 'w') as url_list: base_url = 'https://commons.wikipedia.org' main_page_html = urllib.request.urlopen(f'{base_url}/wiki/Main_Page') soup = BeautifulSoup(main_page_html, 'lxml') # First: Collect ONLY valid Category links (skip None hrefs) category_links = [] for a_tag in soup.find_all('a'): href = a_tag.get('href') # Skip if href is None, and only keep links that contain 'Category' if href and 'Category:' in href: # Convert relative URL to full URL full_url = f'{base_url}{href}' category_links.append(full_url) # Second: For each category page, grab top 10 links for category_url in category_links: try: category_html = urllib.request.urlopen(category_url) category_soup = BeautifulSoup(category_html, 'lxml') # Grab links from the category page (adjust selector if needed) # Wikipedia category pages usually list content in .mw-content-text ul li a page_links = [] for link_tag in category_soup.select('.mw-content-text ul li a'): link_href = link_tag.get('href') if link_href and not link_href.startswith('#'): # Skip anchor links full_page_url = f'{base_url}{link_href}' page_links.append(full_page_url) # Write the first 10 links to file top_10_links = page_links[:10] for link in top_10_links: url_list.write(f'{link}\n') # Optional: Add a separator for readability url_list.write('--- End of Category ---\n') except Exception as e: print(f"Error processing {category_url}: {str(e)}")
Key Improvements:
- Filters out
Nonehrefs: We check ifhrefexists before adding it to our list, which eliminates the TypeError. - Handles relative URLs: Wikipedia uses relative paths, so we append them to the base URL to get full, valid links.
- Safer file handling: The
withstatement automatically closes the file when we're done, preventing resource leaks. - Error handling: A try/except block catches issues like broken links or connection errors, so your scraper doesn't crash unexpectedly.
- Targeted link selection: Uses a CSS selector to grab only relevant content links from category pages, avoiding navigation or anchor links.
内容的提问来源于stack exchange,提问作者Hunter Gary
相关产品推荐
相关产品推荐

