如何用Python验证URL是根域名、子域名还是具体页面?
Great question! Let's walk through two reliable approaches to solve this problem—one using Python's built-in libraries, and another using a third-party package for more accurate handling of complex domain structures.
Approach 1: Using Standard Library (urllib.parse)
This method works well for common domain extensions like .com, .org, .net where the top-level domain (TLD) is a single segment. It avoids installing extra packages.
Step-by-Step Code
from urllib.parse import urlparse def classify_url(url_input): # Add http:// if the URL doesn't have a protocol (prevents parsing errors) if not url_input.startswith(('http://', 'https://')): url_input = f'http://{url_input}' parsed_url = urlparse(url_input) # First check: if there's a non-root path, it's a specific page if parsed_url.path not in ('', '/'): return 'page' # Split the domain into its components domain = parsed_url.netloc domain_parts = domain.split('.') # Classify based on domain structure if len(domain_parts) == 2: # e.g., google.com → root domain return 'root domain' elif len(domain_parts) == 3: # e.g., www.google.com → root; drive.google.com → subdomain return 'root domain' if domain_parts[0] == 'www' else 'sub domain' elif len(domain_parts) > 3: # e.g., mail.drive.google.com → subdomain return 'sub domain' else: return 'invalid domain'
Test the Function
# Test your examples print(classify_url('www.google.com')) # Output: root domain print(classify_url('drive.google.com')) # Output: sub domain print(classify_url('www.google.com/asdasdas')) # Output: page # Additional tests print(classify_url('google.com')) # Output: root domain print(classify_url('mail.drive.google.com'))# Output: sub domain
Limitation
This method struggles with multi-segment TLDs like .co.uk or .com.au. For example, google.co.uk would be incorrectly classified as a subdomain here.
Approach 2: Using tldextract (More Accurate for Complex Domains)
The tldextract package maintains an updated list of TLDs, so it can correctly identify the main domain even for complex extensions.
Step 1: Install the Package
pip install tldextract
Step 2: Code Implementation
import tldextract from urllib.parse import urlparse def classify_url_accurate(url_input): # Add protocol if missing if not url_input.startswith(('http://', 'https://')): url_input = f'http://{url_input}' parsed_url = urlparse(url_input) # Check for specific page first if parsed_url.path not in ('', '/'): return 'page' # Extract domain components accurately extracted = tldextract.extract(url_input) # Classify based on subdomain presence if extracted.subdomain == '' or extracted.subdomain == 'www': # No subdomain or only www → root domain return 'root domain' else: # Any other subdomain → sub domain return 'sub domain'
Test the Function
print(classify_url_accurate('www.google.co.uk')) # Output: root domain print(classify_url_accurate('drive.google.co.uk'))# Output: sub domain print(classify_url_accurate('google.co.uk')) # Output: root domain print(classify_url_accurate('www.google.com/asdasdas')) # Output: page
Key Takeaways
- Use the standard library method for simple, common domains where you don't want extra dependencies.
- Use
tldextractif you need to handle multi-segment TLDs or want more robust, future-proof domain parsing.
内容的提问来源于stack exchange,提问作者Bibin Hashley O P

