You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对JSON推文文件按TEXT字段去重并移除重复文件?

Deduplicate Tweet JSON Files by Unique TEXT Field

Got it, let's tackle this problem step by step. You need to clean up a collection of tweet JSON files, removing any duplicates where the TEXT field content is identical—even if other fields (like user, timestamp, or ID) are different. Here's a practical, safe Python solution that handles this efficiently:

Step 1: Setup & Prerequisites

First, make sure you have a recent version of Python installed. We'll use built-in modules (os, json) plus optional hashlib for memory-efficient duplicate tracking—no extra packages needed.

Step 2: The Script

This script scans your target directory, tracks unique tweet texts, and either moves duplicates to a backup folder or deletes them (backup is recommended to avoid accidental data loss):

import os
import json
import hashlib

def deduplicate_tweet_files(input_dir, backup_dir=None):
    # Create backup directory if specified (safe alternative to deletion)
    if backup_dir and not os.path.exists(backup_dir):
        os.makedirs(backup_dir)
    
    # Track seen tweet content (use hashes for large datasets to save memory)
    seen_content = set()
    
    # Iterate over all files in the input directory
    for filename in os.listdir(input_dir):
        if not filename.endswith(".json"):
            continue  # Skip non-JSON files
        
        file_path = os.path.join(input_dir, filename)
        
        # Handle common file/JSON errors gracefully
        try:
            with open(file_path, "r", encoding="utf-8") as f:
                tweet_data = json.load(f)
            
            # Skip files missing the critical TEXT field
            if "TEXT" not in tweet_data:
                print(f"Skipping {filename}: Missing 'TEXT' field")
                continue
            
            # Optional: Normalize text to catch subtle duplicates (e.g., extra spaces)
            normalized_text = tweet_data["TEXT"].strip()
            # Uncomment below if you want case-insensitive deduplication
            # normalized_text = normalized_text.lower()
            
            # Use hash instead of full text for large datasets (saves memory)
            content_hash = hashlib.sha256(normalized_text.encode("utf-8")).hexdigest()
            
            if content_hash in seen_content:
                # Duplicate found: move to backup or delete
                if backup_dir:
                    dest_path = os.path.join(backup_dir, filename)
                    os.rename(file_path, dest_path)
                    print(f"Moved duplicate {filename} to backup")
                else:
                    os.remove(file_path)
                    print(f"Deleted duplicate {filename}")
            else:
                # Unique content: add to tracking set
                seen_content.add(content_hash)
                print(f"Kept unique file {filename}")
        
        except json.JSONDecodeError:
            print(f"Skipping {filename}: Invalid JSON format")
        except Exception as e:
            print(f"Error processing {filename}: {str(e)}")

# Example usage (replace paths with your actual directories)
if __name__ == "__main__":
    INPUT_DIR = "./your_tweet_json_folder"
    BACKUP_DIR = "./duplicate_backups"  # Set to None to delete duplicates directly
    
    deduplicate_tweet_files(INPUT_DIR, BACKUP_DIR)

Key Customizations & Tips

  • Safety First: Always use the backup_dir parameter for your first run—this lets you review duplicates before permanently deleting anything
  • Normalization: Adjust the normalized_text line to catch more duplicates (e.g., lowercase conversion, removing extra line breaks)
  • Memory Efficiency: Using SHA-256 hashes instead of storing full tweet text is critical if you have thousands of long tweets—this keeps memory usage low
  • Error Handling: The script skips invalid JSON files and those missing the TEXT field, so you don't have to pre-clean your dataset

内容的提问来源于stack exchange,提问作者Bilal Butt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:41:33