IP范围分组及异常检测算法需求:基于运营商IP数据构建有效IP段
Alright, let's break down how to turn your carrier-IP database into actionable valid IP ranges using clustering, so you can spot anomalous new IPs later. Here's a practical, step-by-step approach:
1. Preprocess IPs into Numeric Values
First, we need to convert human-readable IP strings into integers—clustering algorithms work much better with numerical data. For IPv4 addresses, each octet is a byte, so we can shift and combine them into a single 32-bit integer.
Example function in Python:
def ip_to_int(ip_str): octets = list(map(int, ip_str.split('.'))) return octets[0] << 24 | octets[1] << 16 | octets[2] << 8 | octets[3] def int_to_ip(ip_int): return f"{(ip_int >> 24) & 0xFF}.{(ip_int >> 16) & 0xFF}.{(ip_int >> 8) & 0xFF}.{ip_int & 0xFF}"
2. Group IPs by Carrier & Weight by Frequency
Since you care about IP occurrence frequency, start by aggregating your data:
- For each carrier, count how many times each IP appears.
- These counts will act as weights—IPs that show up more often should have more influence on the clusters.
3. Choose the Right Clustering Algorithm
For IP data (which is a 1-dimensional numeric space of integers), DBSCAN is an excellent choice. Here's why:
- It doesn’t require you to predefine the number of clusters (unlike K-Means), which is perfect since you don’t know how many valid ranges each carrier has.
- It automatically identifies dense clusters (valid IP ranges) and marks sparse points as noise (potential outliers).
- You can tune parameters to match your IP segment patterns:
eps: The maximum distance between two IPs to be considered part of the same cluster. For example,eps=256targets C-class subnets (256 consecutive IPs). Adjust this if your carrier uses larger blocks (e.g.,eps=65536for B-class).min_samples: The minimum number of IP instances (weighted by frequency) needed to form a cluster. Set this based on your threshold for "common" IPs (e.g.,min_samples=3means an IP needs to appear at least 3 times to be part of a valid range).
Note: DBSCAN doesn’t natively support weighted points. For large datasets, instead of duplicating rows to simulate weights, look for a weighted DBSCAN implementation or use Density Peak Clustering (DPC), which handles weights natively.
4. Generate Valid IP Ranges from Clusters
Once you’ve clustered the IPs for a carrier:
- For each cluster (excluding noise points), find the minimum and maximum integer values.
- Convert these back to IP strings to get your valid range.
- Optional: Merge adjacent clusters if they form a continuous IP block (e.g., merge
41.74.60.0-41.74.60.255and41.74.61.0-41.74.61.255into41.74.60.0-41.74.61.255).
5. Anomaly Detection Logic
To check if a new IP is anomalous for its carrier:
- Convert the new IP to an integer.
- Check if it falls within any of the carrier’s valid IP ranges.
- If it doesn’t, flag it as an anomaly.
Full Example Code
Here’s a Python snippet tying all this together:
import pandas as pd from sklearn.cluster import DBSCAN import numpy as np # Sample data (replace with your DB query results) df = pd.DataFrame({ 'Name': ['A', 'A', 'A', 'A', 'B', 'B'], 'IP': ['41.74.63.255', '41.74.63.254', '41.74.62.100', '192.168.1.1', '168.167.255.255', '168.167.255.254'] }) # Step 1: Aggregate IP counts per carrier ip_counts = df.groupby(['Name', 'IP']).size().reset_index(name='count') # Step 2: Convert IPs to integers def ip_to_int(ip_str): octets = list(map(int, ip_str.split('.'))) return octets[0] << 24 | octets[1] << 16 | octets[2] << 8 | octets[3] def int_to_ip(ip_int): return f"{(ip_int >> 24) & 0xFF}.{(ip_int >> 16) & 0xFF}.{(ip_int >> 8) & 0xFF}.{ip_int & 0xFF}" ip_counts['ip_int'] = ip_counts['IP'].apply(ip_to_int) # Step 3: Cluster per carrier carrier_ranges = {} for carrier in ip_counts['Name'].unique(): carrier_data = ip_counts[ip_counts['Name'] == carrier] X = carrier_data['ip_int'].values.reshape(-1, 1) weights = carrier_data['count'].values # Simulate weights by repeating rows (good for small datasets) weighted_X = np.repeat(X, weights, axis=0) dbscan = DBSCAN(eps=256, min_samples=2) clusters = dbscan.fit_predict(weighted_X) # Map clusters back to original data carrier_data['cluster'] = np.repeat(clusters, weights) # Generate ranges ranges = [] for cluster_id in set(clusters): if cluster_id == -1: # Skip noise continue cluster_ints = carrier_data[carrier_data['cluster'] == cluster_id]['ip_int'] min_ip = int_to_ip(cluster_ints.min()) max_ip = int_to_ip(cluster_ints.max()) ranges.append((min_ip, max_ip)) carrier_ranges[carrier] = ranges # Step 4: Anomaly check function def is_anomaly(carrier, new_ip): if carrier not in carrier_ranges: return True new_ip_int = ip_to_int(new_ip) for (min_ip, max_ip) in carrier_ranges[carrier]: if ip_to_int(min_ip) <= new_ip_int <= ip_to_int(max_ip): return False return True # Test the function print(is_anomaly('A', '41.74.63.100')) # False (within range) print(is_anomaly('A', '10.0.0.1')) # True (anomaly)
Key Tips
- Tune
epsandmin_samples: Adjust these based on your carrier’s typical IP block sizes and how strict you want the anomaly detection to be. - Handle Large Datasets: For big databases, avoid row duplication—use a weighted clustering library or implement a weighted DBSCAN variant.
- Update Ranges Regularly: IP allocations change over time, so re-run the clustering periodically to keep your valid ranges up to date.
内容的提问来源于stack exchange,提问作者Roni Gadot

