You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python验证URL是根域名、子域名还是具体页面?

Classify URLs into Root Domain, Subdomain, or Specific Page in Python

Great question! Let's walk through two reliable approaches to solve this problem—one using Python's built-in libraries, and another using a third-party package for more accurate handling of complex domain structures.

Approach 1: Using Standard Library (urllib.parse)

This method works well for common domain extensions like .com, .org, .net where the top-level domain (TLD) is a single segment. It avoids installing extra packages.

Step-by-Step Code

from urllib.parse import urlparse

def classify_url(url_input):
    # Add http:// if the URL doesn't have a protocol (prevents parsing errors)
    if not url_input.startswith(('http://', 'https://')):
        url_input = f'http://{url_input}'
    
    parsed_url = urlparse(url_input)
    
    # First check: if there's a non-root path, it's a specific page
    if parsed_url.path not in ('', '/'):
        return 'page'
    
    # Split the domain into its components
    domain = parsed_url.netloc
    domain_parts = domain.split('.')
    
    # Classify based on domain structure
    if len(domain_parts) == 2:
        # e.g., google.com → root domain
        return 'root domain'
    elif len(domain_parts) == 3:
        # e.g., www.google.com → root; drive.google.com → subdomain
        return 'root domain' if domain_parts[0] == 'www' else 'sub domain'
    elif len(domain_parts) > 3:
        # e.g., mail.drive.google.com → subdomain
        return 'sub domain'
    else:
        return 'invalid domain'

Test the Function

# Test your examples
print(classify_url('www.google.com'))       # Output: root domain
print(classify_url('drive.google.com'))     # Output: sub domain
print(classify_url('www.google.com/asdasdas')) # Output: page

# Additional tests
print(classify_url('google.com'))           # Output: root domain
print(classify_url('mail.drive.google.com'))# Output: sub domain

Limitation

This method struggles with multi-segment TLDs like .co.uk or .com.au. For example, google.co.uk would be incorrectly classified as a subdomain here.


Approach 2: Using tldextract (More Accurate for Complex Domains)

The tldextract package maintains an updated list of TLDs, so it can correctly identify the main domain even for complex extensions.

Step 1: Install the Package

pip install tldextract

Step 2: Code Implementation

import tldextract
from urllib.parse import urlparse

def classify_url_accurate(url_input):
    # Add protocol if missing
    if not url_input.startswith(('http://', 'https://')):
        url_input = f'http://{url_input}'
    
    parsed_url = urlparse(url_input)
    
    # Check for specific page first
    if parsed_url.path not in ('', '/'):
        return 'page'
    
    # Extract domain components accurately
    extracted = tldextract.extract(url_input)
    
    # Classify based on subdomain presence
    if extracted.subdomain == '' or extracted.subdomain == 'www':
        # No subdomain or only www → root domain
        return 'root domain'
    else:
        # Any other subdomain → sub domain
        return 'sub domain'

Test the Function

print(classify_url_accurate('www.google.co.uk')) # Output: root domain
print(classify_url_accurate('drive.google.co.uk'))# Output: sub domain
print(classify_url_accurate('google.co.uk'))      # Output: root domain
print(classify_url_accurate('www.google.com/asdasdas')) # Output: page

Key Takeaways

  • Use the standard library method for simple, common domains where you don't want extra dependencies.
  • Use tldextract if you need to handle multi-segment TLDs or want more robust, future-proof domain parsing.

内容的提问来源于stack exchange,提问作者Bibin Hashley O P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:53:42