如何用Python、BeautifulSoup统计网页中同域名指向的链接数量?
Count Same-Domain Links with Python & BeautifulSoup
Got it, you're already on the right path by fetching all links first—let's refine that approach to handle edge cases like relative URLs, domain variations (www vs non-www), and non-HTTP links so you get an accurate count.
Step 1: Install Required Tools
First, make sure you have the necessary packages installed:
pip install requests beautifulsoup4
Step 2: Full Solution Code
Here's a robust script that handles most common scenarios, with comments explaining each part:
import requests from bs4 import BeautifulSoup from urllib.parse import urlparse, urljoin def count_same_domain_links(base_url): # Fetch the webpage (handle request errors gracefully) try: # Add a User-Agent to avoid being blocked by some sites headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"} response = requests.get(base_url, headers=headers) response.raise_for_status() # Raise an error if request fails (4xx/5xx) except requests.exceptions.RequestException as e: print(f"Failed to fetch the page: {str(e)}") return 0 # Parse the HTML with BeautifulSoup soup = BeautifulSoup(response.text, "html.parser") # Extract the base domain (e.g., "www.example.com" from "https://www.example.com/page1") base_domain = urlparse(base_url).netloc # Optional: Ignore "www." prefix to treat www.example.com and example.com as the same # base_domain = base_domain.lstrip("www.") same_domain_count = 0 # Iterate over all <a> tags with an href attribute for a_tag in soup.find_all("a", href=True): href = a_tag["href"] # Skip anchor links (#section), mailto, tel, and other non-web links if href.startswith("#") or href.startswith(("mailto:", "tel:", "javascript:")): continue # Convert relative URLs (like "/about") to absolute URLs (like "https://example.com/about") absolute_url = urljoin(base_url, href) # Extract the domain from the absolute URL link_domain = urlparse(absolute_url).netloc # Optional: Apply the same www-stripping as the base domain # link_domain = link_domain.lstrip("www.") # Check if the link's domain matches the base domain if link_domain == base_domain: same_domain_count += 1 # Optional: Print each matching link to verify # print(f"Matching link: {absolute_url}") return same_domain_count # Example usage if __name__ == "__main__": target_url = "https://www.example.com" # Replace with your target URL total = count_same_domain_links(target_url) print(f"Total same-domain links: {total}")
Key Details to Note
- Relative URL Handling: Using
urljoinensures links like/contactor../blogget converted to full absolute URLs, so we can properly parse their domains. - Error Handling: The script catches request errors (like broken links, timeouts) so it doesn't crash unexpectedly.
- Non-Web Links: We skip anchors, mailto, and JavaScript links since they don't point to other pages on the domain.
- www vs Non-www: Uncomment the
lstrip("www.")lines if you want to treatwww.example.comandexample.comas the same domain.
Advanced: Count Subdomains as Same Domain
If you want to count links to subdomains (e.g., blog.example.com counts as part of example.com), use the tldextract package to get the root registered domain:
- Install it first:
pip install tldextract
- Modify the domain extraction part:
import tldextract def get_root_domain(url): extracted = tldextract.extract(url) return f"{extracted.domain}.{extracted.suffix}" # Then in your function: base_domain = get_root_domain(base_url) link_domain = get_root_domain(absolute_url)
Heads Up!
- Some websites block scrapers without a valid
User-Agentheader—we added one in the script to avoid this. - If the page loads links dynamically with JavaScript, BeautifulSoup won't see them. For those cases, you'll need tools like Selenium or Playwright to render the page first.
内容的提问来源于stack exchange,提问作者Tanmay Maheshwari
相关产品推荐
相关产品推荐

