如何用Python从URL中提取特征?钓鱼网站分类项目技术咨询
Hey there! Since you're already knee-deep in your phishing website detection project—having trained models like Logistic Regression, ANN, and SVM with the UCI dataset—let's break down how to extract actionable features from URLs using Python. These features will add more signal to your models, helping them better spot malicious sites.
Common URL Features to Target
First, let's outline the key features that often distinguish phishing URLs from legitimate ones:
- Basic URL metrics: Total length, domain length
- Suspicious indicators: Presence of IP addresses instead of domain names, special characters like
@or excessive-, percent-encoded content - Domain-related signals: Use of new/rare top-level domains (TLDs), impersonation of well-known brands (e.g.,
fake-google.com) - Path & query parameters: Depth of the URL path, presence/length of query strings (often used to hide malicious payloads)
Python Implementation
Let's build a set of functions to extract these features. We'll use libraries like re for regex, tldextract for domain parsing, and urllib.parse for URL breakdown.
First, install the required packages if you haven't already:
pip install tldextract
Step 1: Import Required Libraries
import re import pandas as pd import tldextract from urllib.parse import urlparse
Step 2: Define Feature Extraction Functions
Let's create modular functions for each category of features:
Basic URL Length
def get_url_length(url): """Return total length of the URL""" return len(url)
Check for IP Address in URL
Phishing sites sometimes use IP addresses instead of legitimate domain names to avoid detection:
def has_ip_address(url): """Check if URL contains an IPv4 address""" ip_regex = re.compile(r'(([01]?\d\d?|2[0-4]\d|25[0-5])\.){3}([01]?\d\d?|2[0-4]\d|25[0-5])') return 1 if ip_regex.search(url) else 0
Detect Suspicious Special Characters
Characters like @ can redirect users to hidden domains, while excessive - or % often signal obfuscation:
def extract_special_char_features(url): """Extract features related to suspicious special characters""" features = { 'has_at': 1 if '@' in url else 0, 'excessive_dashes': 1 if url.count('-') > 3 else 0, 'has_percent_encoding': 1 if '%' in url else 0 } return features
Domain-Related Features
Analyze the domain to spot impersonation or suspicious TLDs:
def extract_domain_features(url): """Extract features from the URL's domain""" extracted = tldextract.extract(url) domain = extracted.domain.lower() tld = extracted.suffix.lower() # Common brand keywords phishers impersonate brand_keywords = ['google', 'facebook', 'amazon', 'paypal', 'bank', 'apple'] return { 'domain_length': len(domain), 'is_new_tld': 1 if tld in ['xyz', 'top', 'club', 'online', 'site'] else 0, 'impersonates_brand': 1 if any(keyword in domain for keyword in brand_keywords) else 0 }
Path & Query Parameter Features
Dig into the URL's path and query strings to find red flags:
def extract_path_query_features(url): """Extract features from URL path and query parameters""" parsed = urlparse(url) path = parsed.path query = parsed.query return { 'path_depth': path.count('/') if path else 0, 'has_query_params': 1 if len(query) > 0 else 0, 'query_length': len(query) }
Step 3: Combine All Features into a Single Function
Wrap all the above functions into one to extract all features for a URL at once:
def extract_all_url_features(url): """Extract all features from a single URL""" features = {} # Add basic length feature features['url_length'] = get_url_length(url) # Add IP address feature features['has_ip'] = has_ip_address(url) # Add special character features features.update(extract_special_char_features(url)) # Add domain features features.update(extract_domain_features(url)) # Add path/query features features.update(extract_path_query_features(url)) return pd.Series(features)
Step 4: Test the Feature Extractor
Let's test this with sample URLs to see how it works:
# Sample URLs (mix of legitimate and phishing) test_urls = [ "https://www.google.com/search?q=phishing+detection", "http://192.168.1.1/fake-bank-login", "https://fake-paypal.xyz/login@malicious-site.com", "https://www.bankofamerica.com/online-banking" ] # Extract features and create a DataFrame features_df = pd.DataFrame([extract_all_url_features(url) for url in test_urls], index=test_urls) print(features_df)
Integrating with Your Existing Project
Once you've extracted these features, you can merge them with your UCI dataset (make sure each row corresponds to the correct URL) and use the combined feature set to retrain your models. You might find that adding these URL-specific features boosts your model's accuracy!
You can also expand this further by adding features like:
- Whether the URL uses HTTPS (phishing sites often use HTTP or invalid SSL certificates)
- Domain registration age (using the
python-whoislibrary) - Number of subdomains (phishers often use long, nested subdomains)
内容的提问来源于stack exchange,提问作者Abhishek Dobhal

