You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于CSV/TXT中的IP与时间戳批量过滤Shodan JSON Banner

Got it, let's solve this batch filtering problem efficiently—no more manual jq commands for each pair! Here are two solid approaches you can use depending on your comfort with shell scripting vs. Python:

Approach 1: Shell + jq (Fast, no extra dependencies)

This leverages jq's built-in functions to load your match pairs and filter the Shodan JSON in one go.

Step 1: Prepare your match file

First, make sure your CSV/TXT has each ip_str and timestamp pair on a single line. For a CSV with a header (e.g., ip_str,timestamp), we'll skip the header line. For a plain TXT file with no header, remove the sed 1d part from the command below.

Step 2: Run the combined command

# 1. Convert your match pairs into a JSON array
matches=$(jq -nR '[inputs | split(",") | {"ip_str": .[0], "timestamp": .[1]}]' <(sed 1d matches.csv))

# 2. Filter the Shodan JSON against the match array
jq --argjson matches "$matches" '.[] | select(. as $item | $matches | any(.ip_str == $item.ip_str and .timestamp == $item.timestamp))' extract_3month_fromshodan.json > filtered_results.json

How this works:

  • sed 1d matches.csv: Skips the first header line (remove this if your file has no header).
  • The first jq command turns each line of your match file into a JSON object like {"ip_str":"192.168.1.1","timestamp":"2024-05-20T10:00:00Z"}, then wraps all of them into an array.
  • --argjson matches "$matches" passes that array into the filtering jq command.
  • any(...) checks if the current Shodan entry matches any of the ip/timestamp pairs in your match list.

Approach 2: Python Script (Flexible, easy to customize)

If you're more comfortable with Python, or need to handle edge cases like timestamp format mismatches, this is a great alternative.

Script Code

import json
import csv

# --------------------------
# Configure your files here
MATCH_FILE = "matches.csv"  # or "matches.txt"
SHODAN_FILE = "extract_3month_fromshodan.json"
OUTPUT_FILE = "filtered_results.json"
# --------------------------

# Load match pairs into a set for fast lookup
match_set = set()

# Option A: For CSV with headers (ip_str, timestamp)
with open(MATCH_FILE, "r") as f:
    reader = csv.DictReader(f)
    for row in reader:
        # Store as a tuple (ip, timestamp)
        match_set.add((row["ip_str"].strip(), row["timestamp"].strip()))

# Option B: For plain TXT (each line is "ip,timestamp" with no header)
# with open(MATCH_FILE, "r") as f:
#     for line in f:
#         line = line.strip()
#         if not line:
#             continue
#         # Split only once in case timestamp has commas
#         ip, timestamp = line.split(",", 1)
#         match_set.add((ip.strip(), timestamp.strip()))

# Filter the Shodan data
filtered_results = []

# Case 1: Shodan file is JSON Lines (one banner per line)
with open(SHODAN_FILE, "r") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        banner = json.loads(line)
        current_pair = (banner.get("ip_str"), banner.get("timestamp"))
        if current_pair in match_set:
            filtered_results.append(banner)

# Case 2: Shodan file is a single JSON array (uncomment if needed)
# with open(SHODAN_FILE, "r") as f:
#     all_banners = json.load(f)
#     for banner in all_banners:
#         current_pair = (banner.get("ip_str"), banner.get("timestamp"))
#         if current_pair in match_set:
#             filtered_results.append(banner)

# Save the filtered results
with open(OUTPUT_FILE, "w") as f:
    json.dump(filtered_results, f, indent=2)

How this works:

  • We store your match pairs in a set for O(1) lookup speed, which is efficient even with 290 pairs.
  • The script reads the Shodan file either as JSON Lines (common for Shodan exports) or a single JSON array—just uncomment the relevant section.
  • It checks each banner's ip_str and timestamp against the match set, and collects all matching entries into the output.

Key Notes for Both Approaches

  • Timestamp Format: Ensure the timestamp values in your match file exactly match the format in the Shodan JSON (e.g., 2024-05-20T12:34:56.789Z). If there are format differences, you'll need to add parsing logic (e.g., using Python's datetime module to normalize timestamps).
  • Large Files: Both approaches handle large files efficiently—jq streams the input, and the Python script reads line-by-line instead of loading the entire file into memory.

内容的提问来源于stack exchange,提问作者Taskin Alam Jim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:47:40