You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从URL中提取特征?钓鱼网站分类项目技术咨询

Extracting URL Features for Phishing Detection in Python

Hey there! Since you're already knee-deep in your phishing website detection project—having trained models like Logistic Regression, ANN, and SVM with the UCI dataset—let's break down how to extract actionable features from URLs using Python. These features will add more signal to your models, helping them better spot malicious sites.

Common URL Features to Target

First, let's outline the key features that often distinguish phishing URLs from legitimate ones:

  • Basic URL metrics: Total length, domain length
  • Suspicious indicators: Presence of IP addresses instead of domain names, special characters like @ or excessive -, percent-encoded content
  • Domain-related signals: Use of new/rare top-level domains (TLDs), impersonation of well-known brands (e.g., fake-google.com)
  • Path & query parameters: Depth of the URL path, presence/length of query strings (often used to hide malicious payloads)

Python Implementation

Let's build a set of functions to extract these features. We'll use libraries like re for regex, tldextract for domain parsing, and urllib.parse for URL breakdown.

First, install the required packages if you haven't already:

pip install tldextract

Step 1: Import Required Libraries

import re
import pandas as pd
import tldextract
from urllib.parse import urlparse

Step 2: Define Feature Extraction Functions

Let's create modular functions for each category of features:

Basic URL Length

def get_url_length(url):
    """Return total length of the URL"""
    return len(url)

Check for IP Address in URL

Phishing sites sometimes use IP addresses instead of legitimate domain names to avoid detection:

def has_ip_address(url):
    """Check if URL contains an IPv4 address"""
    ip_regex = re.compile(r'(([01]?\d\d?|2[0-4]\d|25[0-5])\.){3}([01]?\d\d?|2[0-4]\d|25[0-5])')
    return 1 if ip_regex.search(url) else 0

Detect Suspicious Special Characters

Characters like @ can redirect users to hidden domains, while excessive - or % often signal obfuscation:

def extract_special_char_features(url):
    """Extract features related to suspicious special characters"""
    features = {
        'has_at': 1 if '@' in url else 0,
        'excessive_dashes': 1 if url.count('-') > 3 else 0,
        'has_percent_encoding': 1 if '%' in url else 0
    }
    return features

Domain-Related Features

Analyze the domain to spot impersonation or suspicious TLDs:

def extract_domain_features(url):
    """Extract features from the URL's domain"""
    extracted = tldextract.extract(url)
    domain = extracted.domain.lower()
    tld = extracted.suffix.lower()
    
    # Common brand keywords phishers impersonate
    brand_keywords = ['google', 'facebook', 'amazon', 'paypal', 'bank', 'apple']
    
    return {
        'domain_length': len(domain),
        'is_new_tld': 1 if tld in ['xyz', 'top', 'club', 'online', 'site'] else 0,
        'impersonates_brand': 1 if any(keyword in domain for keyword in brand_keywords) else 0
    }

Path & Query Parameter Features

Dig into the URL's path and query strings to find red flags:

def extract_path_query_features(url):
    """Extract features from URL path and query parameters"""
    parsed = urlparse(url)
    path = parsed.path
    query = parsed.query
    
    return {
        'path_depth': path.count('/') if path else 0,
        'has_query_params': 1 if len(query) > 0 else 0,
        'query_length': len(query)
    }

Step 3: Combine All Features into a Single Function

Wrap all the above functions into one to extract all features for a URL at once:

def extract_all_url_features(url):
    """Extract all features from a single URL"""
    features = {}
    
    # Add basic length feature
    features['url_length'] = get_url_length(url)
    
    # Add IP address feature
    features['has_ip'] = has_ip_address(url)
    
    # Add special character features
    features.update(extract_special_char_features(url))
    
    # Add domain features
    features.update(extract_domain_features(url))
    
    # Add path/query features
    features.update(extract_path_query_features(url))
    
    return pd.Series(features)

Step 4: Test the Feature Extractor

Let's test this with sample URLs to see how it works:

# Sample URLs (mix of legitimate and phishing)
test_urls = [
    "https://www.google.com/search?q=phishing+detection",
    "http://192.168.1.1/fake-bank-login",
    "https://fake-paypal.xyz/login@malicious-site.com",
    "https://www.bankofamerica.com/online-banking"
]

# Extract features and create a DataFrame
features_df = pd.DataFrame([extract_all_url_features(url) for url in test_urls], index=test_urls)
print(features_df)

Integrating with Your Existing Project

Once you've extracted these features, you can merge them with your UCI dataset (make sure each row corresponds to the correct URL) and use the combined feature set to retrain your models. You might find that adding these URL-specific features boosts your model's accuracy!

You can also expand this further by adding features like:

  • Whether the URL uses HTTPS (phishing sites often use HTTP or invalid SSL certificates)
  • Domain registration age (using the python-whois library)
  • Number of subdomains (phishers often use long, nested subdomains)

内容的提问来源于stack exchange,提问作者Abhishek Dobhal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:49:38