You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python、BeautifulSoup统计网页中同域名指向的链接数量?

Got it, you're already on the right path by fetching all links first—let's refine that approach to handle edge cases like relative URLs, domain variations (www vs non-www), and non-HTTP links so you get an accurate count.

Step 1: Install Required Tools

First, make sure you have the necessary packages installed:

pip install requests beautifulsoup4

Step 2: Full Solution Code

Here's a robust script that handles most common scenarios, with comments explaining each part:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urlparse, urljoin

def count_same_domain_links(base_url):
    # Fetch the webpage (handle request errors gracefully)
    try:
        # Add a User-Agent to avoid being blocked by some sites
        headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
        response = requests.get(base_url, headers=headers)
        response.raise_for_status()  # Raise an error if request fails (4xx/5xx)
    except requests.exceptions.RequestException as e:
        print(f"Failed to fetch the page: {str(e)}")
        return 0

    # Parse the HTML with BeautifulSoup
    soup = BeautifulSoup(response.text, "html.parser")

    # Extract the base domain (e.g., "www.example.com" from "https://www.example.com/page1")
    base_domain = urlparse(base_url).netloc
    # Optional: Ignore "www." prefix to treat www.example.com and example.com as the same
    # base_domain = base_domain.lstrip("www.")

    same_domain_count = 0

    # Iterate over all <a> tags with an href attribute
    for a_tag in soup.find_all("a", href=True):
        href = a_tag["href"]
        
        # Skip anchor links (#section), mailto, tel, and other non-web links
        if href.startswith("#") or href.startswith(("mailto:", "tel:", "javascript:")):
            continue

        # Convert relative URLs (like "/about") to absolute URLs (like "https://example.com/about")
        absolute_url = urljoin(base_url, href)
        
        # Extract the domain from the absolute URL
        link_domain = urlparse(absolute_url).netloc
        # Optional: Apply the same www-stripping as the base domain
        # link_domain = link_domain.lstrip("www.")

        # Check if the link's domain matches the base domain
        if link_domain == base_domain:
            same_domain_count += 1
            # Optional: Print each matching link to verify
            # print(f"Matching link: {absolute_url}")

    return same_domain_count

# Example usage
if __name__ == "__main__":
    target_url = "https://www.example.com"  # Replace with your target URL
    total = count_same_domain_links(target_url)
    print(f"Total same-domain links: {total}")

Key Details to Note

  • Relative URL Handling: Using urljoin ensures links like /contact or ../blog get converted to full absolute URLs, so we can properly parse their domains.
  • Error Handling: The script catches request errors (like broken links, timeouts) so it doesn't crash unexpectedly.
  • Non-Web Links: We skip anchors, mailto, and JavaScript links since they don't point to other pages on the domain.
  • www vs Non-www: Uncomment the lstrip("www.") lines if you want to treat www.example.com and example.com as the same domain.

Advanced: Count Subdomains as Same Domain

If you want to count links to subdomains (e.g., blog.example.com counts as part of example.com), use the tldextract package to get the root registered domain:

  1. Install it first:
pip install tldextract
  1. Modify the domain extraction part:
import tldextract

def get_root_domain(url):
    extracted = tldextract.extract(url)
    return f"{extracted.domain}.{extracted.suffix}"

# Then in your function:
base_domain = get_root_domain(base_url)
link_domain = get_root_domain(absolute_url)

Heads Up!

  • Some websites block scrapers without a valid User-Agent header—we added one in the script to avoid this.
  • If the page loads links dynamically with JavaScript, BeautifulSoup won't see them. For those cases, you'll need tools like Selenium or Playwright to render the page first.

内容的提问来源于stack exchange,提问作者Tanmay Maheshwari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:00:48