You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3如何在指定域名下进行网页爬取与内容抓取?

Fixing Domain Restriction for Your Web Crawler

Great question—your current approach is on the right track, but the exact hostname match is likely tripping you up (think subdomains like www.newyorktimes.com or mobile.newyorktimes.com not matching your ["newyorktimes.com"] list). Let’s break down practical fixes and best practices to keep your crawler locked to your target domains.

Key Issues with Your Current Code

Your check p.hostname in domains only works for exact hostname matches. It fails for:

  • Subdomains (e.g., www.newyorktimes.com vs newyorktimes.com)
  • Relative URLs (like /section/politics) that don’t have a hostname yet
  • Edge cases like misformatted URLs or non-HTTP schemes (e.g., mailto: links)

Step-by-Step Solutions

1. Resolve Relative URLs First

Before checking the domain, convert any relative links to absolute URLs using the base URL of the page you’re crawling. This ensures urlparse can extract a valid hostname:

from urllib.parse import urljoin, urlparse

base_url = "https://www.newyorktimes.com"
relative_link = "/section/politics"
absolute_url = urljoin(base_url, relative_link)  # Returns "https://www.newyorktimes.com/section/politics"

2. Check for Domain (and Subdomain) Match

Instead of exact string matching, verify if the hostname ends with your target domain (add a leading dot to avoid false matches like newyorktimes.com.co). Here’s a reusable function:

def is_allowed_domain(url, allowed_domains):
    parsed = urlparse(url)
    # Skip URLs without a hostname (e.g., mailto:, javascript:)
    if not parsed.hostname or parsed.scheme not in ["http", "https"]:
        return False
    for domain in allowed_domains:
        # Match exact domain OR subdomain of the target
        if parsed.hostname == domain or parsed.hostname.endswith(f".{domain}"):
            return True
    return False

Use it like this:

allowed_domains = ["newyorktimes.com"]
url = "https://mobile.newyorktimes.com/tech"
if is_allowed_domain(url, allowed_domains):
    # Process the URL and extract content
    pass
else:
    return []

3. Leverage Scrapy’s Built-in Tools (If Using Scrapy)

Since you mentioned Scrapy, it handles domain restriction out of the box—no need to reinvent the wheel. Just set the allowed_domains attribute in your spider:

import scrapy

class NYTimesSpider(scrapy.Spider):
    name = "nytimes_crawler"
    allowed_domains = ["newyorktimes.com"]
    start_urls = ["https://www.newyorktimes.com"]

    def parse(self, response):
        # Extract text content from the page
        page_text = " ".join(response.xpath('//body//text()').getall()).strip()
        
        # Follow all valid links within the allowed domain
        for link in response.css("a::attr(href)"):
            yield response.follow(link, self.parse)

Scrapy automatically resolves relative URLs, skips non-HTTP links, and blocks requests outside your allowed_domains.

4. Extra Tips to Polish Your Crawler

  • Normalize URLs: Remove fragments (the part after #) to avoid crawling the same page multiple times:
    normalized_url = parsed._replace(fragment="").geturl()
    
  • Track Visited URLs: Use a set to store already crawled URLs and avoid duplicates:
    visited_urls = set()
    if normalized_url not in visited_urls:
        visited_urls.add(normalized_url)
        # Process the URL
    
  • Respect robots.txt: Use robotparser or Scrapy’s built-in compliance to avoid crawling restricted paths.

内容的提问来源于stack exchange,提问作者Anon Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:32:35