You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

IP范围分组及异常检测算法需求:基于运营商IP数据构建有效IP段

Clustering IPs by Carrier to Build Valid Ranges for Anomaly Detection

Alright, let's break down how to turn your carrier-IP database into actionable valid IP ranges using clustering, so you can spot anomalous new IPs later. Here's a practical, step-by-step approach:

1. Preprocess IPs into Numeric Values

First, we need to convert human-readable IP strings into integers—clustering algorithms work much better with numerical data. For IPv4 addresses, each octet is a byte, so we can shift and combine them into a single 32-bit integer.

Example function in Python:

def ip_to_int(ip_str):
    octets = list(map(int, ip_str.split('.')))
    return octets[0] << 24 | octets[1] << 16 | octets[2] << 8 | octets[3]

def int_to_ip(ip_int):
    return f"{(ip_int >> 24) & 0xFF}.{(ip_int >> 16) & 0xFF}.{(ip_int >> 8) & 0xFF}.{ip_int & 0xFF}"

2. Group IPs by Carrier & Weight by Frequency

Since you care about IP occurrence frequency, start by aggregating your data:

  • For each carrier, count how many times each IP appears.
  • These counts will act as weights—IPs that show up more often should have more influence on the clusters.

3. Choose the Right Clustering Algorithm

For IP data (which is a 1-dimensional numeric space of integers), DBSCAN is an excellent choice. Here's why:

  • It doesn’t require you to predefine the number of clusters (unlike K-Means), which is perfect since you don’t know how many valid ranges each carrier has.
  • It automatically identifies dense clusters (valid IP ranges) and marks sparse points as noise (potential outliers).
  • You can tune parameters to match your IP segment patterns:
    • eps: The maximum distance between two IPs to be considered part of the same cluster. For example, eps=256 targets C-class subnets (256 consecutive IPs). Adjust this if your carrier uses larger blocks (e.g., eps=65536 for B-class).
    • min_samples: The minimum number of IP instances (weighted by frequency) needed to form a cluster. Set this based on your threshold for "common" IPs (e.g., min_samples=3 means an IP needs to appear at least 3 times to be part of a valid range).

Note: DBSCAN doesn’t natively support weighted points. For large datasets, instead of duplicating rows to simulate weights, look for a weighted DBSCAN implementation or use Density Peak Clustering (DPC), which handles weights natively.

4. Generate Valid IP Ranges from Clusters

Once you’ve clustered the IPs for a carrier:

  • For each cluster (excluding noise points), find the minimum and maximum integer values.
  • Convert these back to IP strings to get your valid range.
  • Optional: Merge adjacent clusters if they form a continuous IP block (e.g., merge 41.74.60.0-41.74.60.255 and 41.74.61.0-41.74.61.255 into 41.74.60.0-41.74.61.255).

5. Anomaly Detection Logic

To check if a new IP is anomalous for its carrier:

  1. Convert the new IP to an integer.
  2. Check if it falls within any of the carrier’s valid IP ranges.
  3. If it doesn’t, flag it as an anomaly.

Full Example Code

Here’s a Python snippet tying all this together:

import pandas as pd
from sklearn.cluster import DBSCAN
import numpy as np

# Sample data (replace with your DB query results)
df = pd.DataFrame({
    'Name': ['A', 'A', 'A', 'A', 'B', 'B'],
    'IP': ['41.74.63.255', '41.74.63.254', '41.74.62.100', '192.168.1.1', '168.167.255.255', '168.167.255.254']
})

# Step 1: Aggregate IP counts per carrier
ip_counts = df.groupby(['Name', 'IP']).size().reset_index(name='count')

# Step 2: Convert IPs to integers
def ip_to_int(ip_str):
    octets = list(map(int, ip_str.split('.')))
    return octets[0] << 24 | octets[1] << 16 | octets[2] << 8 | octets[3]

def int_to_ip(ip_int):
    return f"{(ip_int >> 24) & 0xFF}.{(ip_int >> 16) & 0xFF}.{(ip_int >> 8) & 0xFF}.{ip_int & 0xFF}"

ip_counts['ip_int'] = ip_counts['IP'].apply(ip_to_int)

# Step 3: Cluster per carrier
carrier_ranges = {}
for carrier in ip_counts['Name'].unique():
    carrier_data = ip_counts[ip_counts['Name'] == carrier]
    X = carrier_data['ip_int'].values.reshape(-1, 1)
    weights = carrier_data['count'].values
    
    # Simulate weights by repeating rows (good for small datasets)
    weighted_X = np.repeat(X, weights, axis=0)
    dbscan = DBSCAN(eps=256, min_samples=2)
    clusters = dbscan.fit_predict(weighted_X)
    
    # Map clusters back to original data
    carrier_data['cluster'] = np.repeat(clusters, weights)
    
    # Generate ranges
    ranges = []
    for cluster_id in set(clusters):
        if cluster_id == -1:  # Skip noise
            continue
        cluster_ints = carrier_data[carrier_data['cluster'] == cluster_id]['ip_int']
        min_ip = int_to_ip(cluster_ints.min())
        max_ip = int_to_ip(cluster_ints.max())
        ranges.append((min_ip, max_ip))
    
    carrier_ranges[carrier] = ranges

# Step 4: Anomaly check function
def is_anomaly(carrier, new_ip):
    if carrier not in carrier_ranges:
        return True
    new_ip_int = ip_to_int(new_ip)
    for (min_ip, max_ip) in carrier_ranges[carrier]:
        if ip_to_int(min_ip) <= new_ip_int <= ip_to_int(max_ip):
            return False
    return True

# Test the function
print(is_anomaly('A', '41.74.63.100'))  # False (within range)
print(is_anomaly('A', '10.0.0.1'))      # True (anomaly)

Key Tips

  • Tune eps and min_samples: Adjust these based on your carrier’s typical IP block sizes and how strict you want the anomaly detection to be.
  • Handle Large Datasets: For big databases, avoid row duplication—use a weighted clustering library or implement a weighted DBSCAN variant.
  • Update Ranges Regularly: IP allocations change over time, so re-run the clustering periodically to keep your valid ranges up to date.

内容的提问来源于stack exchange,提问作者Roni Gadot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:40:07