Python 3如何在指定域名下进行网页爬取与内容抓取?
Great question—your current approach is on the right track, but the exact hostname match is likely tripping you up (think subdomains like www.newyorktimes.com or mobile.newyorktimes.com not matching your ["newyorktimes.com"] list). Let’s break down practical fixes and best practices to keep your crawler locked to your target domains.
Key Issues with Your Current Code
Your check p.hostname in domains only works for exact hostname matches. It fails for:
- Subdomains (e.g.,
www.newyorktimes.comvsnewyorktimes.com) - Relative URLs (like
/section/politics) that don’t have a hostname yet - Edge cases like misformatted URLs or non-HTTP schemes (e.g.,
mailto:links)
Step-by-Step Solutions
1. Resolve Relative URLs First
Before checking the domain, convert any relative links to absolute URLs using the base URL of the page you’re crawling. This ensures urlparse can extract a valid hostname:
from urllib.parse import urljoin, urlparse base_url = "https://www.newyorktimes.com" relative_link = "/section/politics" absolute_url = urljoin(base_url, relative_link) # Returns "https://www.newyorktimes.com/section/politics"
2. Check for Domain (and Subdomain) Match
Instead of exact string matching, verify if the hostname ends with your target domain (add a leading dot to avoid false matches like newyorktimes.com.co). Here’s a reusable function:
def is_allowed_domain(url, allowed_domains): parsed = urlparse(url) # Skip URLs without a hostname (e.g., mailto:, javascript:) if not parsed.hostname or parsed.scheme not in ["http", "https"]: return False for domain in allowed_domains: # Match exact domain OR subdomain of the target if parsed.hostname == domain or parsed.hostname.endswith(f".{domain}"): return True return False
Use it like this:
allowed_domains = ["newyorktimes.com"] url = "https://mobile.newyorktimes.com/tech" if is_allowed_domain(url, allowed_domains): # Process the URL and extract content pass else: return []
3. Leverage Scrapy’s Built-in Tools (If Using Scrapy)
Since you mentioned Scrapy, it handles domain restriction out of the box—no need to reinvent the wheel. Just set the allowed_domains attribute in your spider:
import scrapy class NYTimesSpider(scrapy.Spider): name = "nytimes_crawler" allowed_domains = ["newyorktimes.com"] start_urls = ["https://www.newyorktimes.com"] def parse(self, response): # Extract text content from the page page_text = " ".join(response.xpath('//body//text()').getall()).strip() # Follow all valid links within the allowed domain for link in response.css("a::attr(href)"): yield response.follow(link, self.parse)
Scrapy automatically resolves relative URLs, skips non-HTTP links, and blocks requests outside your allowed_domains.
4. Extra Tips to Polish Your Crawler
- Normalize URLs: Remove fragments (the part after
#) to avoid crawling the same page multiple times:normalized_url = parsed._replace(fragment="").geturl() - Track Visited URLs: Use a set to store already crawled URLs and avoid duplicates:
visited_urls = set() if normalized_url not in visited_urls: visited_urls.add(normalized_url) # Process the URL - Respect
robots.txt: Userobotparseror Scrapy’s built-in compliance to avoid crawling restricted paths.
内容的提问来源于stack exchange,提问作者Anon Li

