You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按日期、指定端口57621及源IP分组聚合统计

Solution: Date-wise Aggregation for Specific Port & Source IP

Let's break down how to adjust your code to meet the exact requirements: filter for port 57621, group by date (from your Time index), and calculate multi-dimensional packet length stats per source IP.

Step 1: Filter and Aggregate the Data

First, we'll narrow down the data to only rows with your target port, then group by date (extracted from the datetime index) and source IP. We'll compute all the aggregation metrics you need in one go:

import pandas as pd
import matplotlib.pyplot as plt

# Assuming your dataframe is named 'df' with a datetime index (Time column)
# 1. Keep only rows where destination port is 57621
filtered_df = df[df['Dst_Port'] == 57621]

# 2. Group by date (from index) and Source_IP, then aggregate Packet_length
aggregated_results = filtered_df.groupby([filtered_df.index.date, 'Source_IP'])['Packet_length'].agg(
    total_length='sum',
    min_length='min',
    max_length='max',
    avg_length='mean',
    median_length='median',
    std_dev='std'
).reset_index()

# Rename the auto-generated date column for clarity
aggregated_results.rename(columns={'level_0': 'Date'}, inplace=True)

# Print the results to verify
print(aggregated_results)

This will output a dataframe where each row represents a unique date + source IP pair, with all your requested packet length metrics.

Step 2: Visualize the Aggregated Data

To make the results easy to interpret, let's create subplots for key metrics (total, average, max length) grouped by date and source IP:

# Pivot the data for cleaner plotting
pivoted_total = aggregated_results.pivot(index='Date', columns='Source_IP', values='total_length')
pivoted_avg = aggregated_results.pivot(index='Date', columns='Source_IP', values='avg_length')
pivoted_max = aggregated_results.pivot(index='Date', columns='Source_IP', values='max_length')

# Create a 3-row subplot grid
fig, (ax1, ax2, ax3) = plt.subplots(3, 1, figsize=(10, 16))

# Plot total packet length per date/IP
pivoted_total.plot(kind='bar', ax=ax1, title='Total Packet Length per Date & Source IP')
ax1.set_ylabel('Total Length')
ax1.tick_params(axis='x', rotation=45)

# Plot average packet length per date/IP
pivoted_avg.plot(kind='bar', ax=ax2, title='Average Packet Length per Date & Source IP')
ax2.set_ylabel('Average Length')
ax2.tick_params(axis='x', rotation=45)

# Plot maximum packet length per date/IP
pivoted_max.plot(kind='bar', ax=ax3, title='Maximum Packet Length per Date & Source IP')
ax3.set_ylabel('Maximum Length')
ax3.tick_params(axis='x', rotation=45)

# Adjust layout to prevent label overlap
plt.tight_layout()
plt.show()

Quick Notes:

  • If your Time index isn't already a datetime type, convert it first with df.index = pd.to_datetime(df.index)
  • Add or remove aggregation metrics by modifying the agg() call (e.g., add packet_count='count' to track number of packets)
  • For a more compact view, you could plot all metrics in a single grouped bar chart, but subplots make it easier to compare each metric individually

内容的提问来源于stack exchange,提问作者delwar.naist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:25:06