You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python分析PCAP文件:查找感染前非谷歌搜索引擎的访问信息

Solution to Parse Bing Search Activity from PCAP with Python

Got it, let's fix this up! Your current code works for detecting search engines, but parsing raw PCAP bytes directly isn't reliable for extracting specific web activity details. We'll use Scapy—a robust Python library for packet analysis—to properly parse HTTP/HTTPS traffic and solve all four tasks.

Prerequisite: Install Scapy

First, install the library if you haven't already:

pip install scapy

Step-by-Step Implementation

Here's the complete code with explanations for each task:

from scapy.all import rdpcap, HTTPRequest, HTTPResponse
from urllib.parse import urlparse, parse_qs, unquote
import re

# Load the PCAP file
packets = rdpcap('nameofpcapfile.pcap')

# --------------------------
# 1. Confirm the Search Engine (Optimized)
# --------------------------
search_engine = None
for packet in packets:
    if packet.haslayer(HTTPRequest):
        host = packet[HTTPRequest].Host.decode('utf-8', errors='ignore')
        if 'bing.com' in host:
            search_engine = "Bing (www.bing.com)"
            break

if not search_engine:
    print("Search engine not found in PCAP")
else:
    print(f"Search Engine Used: {search_engine}")

# --------------------------
# 2. Extract Search Keywords
# --------------------------
search_keywords = ""
# Find Bing search requests (targets /search endpoint)
bing_search_requests = []
for packet in packets:
    if packet.haslayer(HTTPRequest):
        host = packet[HTTPRequest].Host.decode('utf-8', errors='ignore')
        path = packet[HTTPRequest].Path.decode('utf-8', errors='ignore')
        if 'bing.com' in host and '/search' in path:
            bing_search_requests.append(packet)

if bing_search_requests:
    # Parse the first search request (adjust if multiple searches exist)
    request_path = bing_search_requests[0][HTTPRequest].Path.decode('utf-8')
    query_params = parse_qs(urlparse(request_path).query)
    # Bing uses 'q' parameter for search keywords
    search_keywords = query_params.get('q', ['No keywords found'])[0]
    print(f"Search Keywords: {search_keywords}")
else:
    print("No Bing search requests found in PCAP")

# --------------------------
# 3. Find Recommended Websites from Bing
# --------------------------
recommended_sites = set()
# Locate Bing's search result HTML responses
bing_response_packets = []
for packet in packets:
    if packet.haslayer(HTTPResponse) and packet.haslayer('Raw'):
        # Check if response is from Bing
        if packet.haslayer(HTTPRequest):
            host = packet[HTTPRequest].Host.decode('utf-8', errors='ignore')
            if 'bing.com' in host:
                # Verify it's an HTML response (search results are HTML)
                content_type = packet[HTTPResponse].get('Content-Type', b'').decode('utf-8', errors='ignore')
                if 'text/html' in content_type:
                    bing_response_packets.append(packet)

# Extract links from HTML content
for resp_packet in bing_response_packets:
    html_content = resp_packet['Raw'].load.decode('utf-8', errors='ignore')
    # Bing uses /url?q=... for result links (to track clicks)
    link_matches = re.findall(r'/url\?q=(https?://[^&]+)', html_content)
    for link in link_matches:
        decoded_link = unquote(link)
        # Extract the domain name from the URL
        domain = urlparse(decoded_link).netloc
        recommended_sites.add(domain)

if recommended_sites:
    print(f"Recommended Sites from Bing: {list(recommended_sites)}")
else:
    print("No recommended sites found in Bing responses")

# --------------------------
# 4. Locate Actually Visited Websites
# --------------------------
visited_sites = set()
if bing_search_requests:
    # Get the timestamp of the first search request
    search_time = bing_search_requests[0].time
    # Look for HTTP requests made after the search (exclude Bing itself)
    for packet in packets:
        if packet.time > search_time and packet.haslayer(HTTPRequest):
            host = packet[HTTPRequest].Host.decode('utf-8', errors='ignore')
            if 'bing.com' not in host:
                visited_sites.add(host)

if visited_sites:
    print(f"Actual Sites Visited After Search: {list(visited_sites)}")
else:
    print("No sites visited after the Bing search found in PCAP")

Key Notes

  • HTTPS Traffic: If the PCAP contains encrypted HTTPS traffic, Scapy can't parse the content by default. You'll need to:
    1. Use Wireshark to decrypt the traffic (using SSL/TLS master secrets) and export a decrypted PCAP, or
    2. Configure Scapy with the necessary SSL keys to decrypt the packets.
  • Multiple Searches: The code assumes one search request. If there are multiple, adjust the logic to loop through all bing_search_requests instead of just the first.
  • Edge Cases: Some packets might have encoding issues, so we use errors='ignore' to avoid crashes—you can refine this if needed for your specific PCAP.

内容的提问来源于stack exchange,提问作者SNIPERATI0N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:39:10