如何对JSON推文文件按TEXT字段去重并移除重复文件?
TEXT Field Got it, let's tackle this problem step by step. You need to clean up a collection of tweet JSON files, removing any duplicates where the TEXT field content is identical—even if other fields (like user, timestamp, or ID) are different. Here's a practical, safe Python solution that handles this efficiently:
Step 1: Setup & Prerequisites
First, make sure you have a recent version of Python installed. We'll use built-in modules (os, json) plus optional hashlib for memory-efficient duplicate tracking—no extra packages needed.
Step 2: The Script
This script scans your target directory, tracks unique tweet texts, and either moves duplicates to a backup folder or deletes them (backup is recommended to avoid accidental data loss):
import os import json import hashlib def deduplicate_tweet_files(input_dir, backup_dir=None): # Create backup directory if specified (safe alternative to deletion) if backup_dir and not os.path.exists(backup_dir): os.makedirs(backup_dir) # Track seen tweet content (use hashes for large datasets to save memory) seen_content = set() # Iterate over all files in the input directory for filename in os.listdir(input_dir): if not filename.endswith(".json"): continue # Skip non-JSON files file_path = os.path.join(input_dir, filename) # Handle common file/JSON errors gracefully try: with open(file_path, "r", encoding="utf-8") as f: tweet_data = json.load(f) # Skip files missing the critical TEXT field if "TEXT" not in tweet_data: print(f"Skipping {filename}: Missing 'TEXT' field") continue # Optional: Normalize text to catch subtle duplicates (e.g., extra spaces) normalized_text = tweet_data["TEXT"].strip() # Uncomment below if you want case-insensitive deduplication # normalized_text = normalized_text.lower() # Use hash instead of full text for large datasets (saves memory) content_hash = hashlib.sha256(normalized_text.encode("utf-8")).hexdigest() if content_hash in seen_content: # Duplicate found: move to backup or delete if backup_dir: dest_path = os.path.join(backup_dir, filename) os.rename(file_path, dest_path) print(f"Moved duplicate {filename} to backup") else: os.remove(file_path) print(f"Deleted duplicate {filename}") else: # Unique content: add to tracking set seen_content.add(content_hash) print(f"Kept unique file {filename}") except json.JSONDecodeError: print(f"Skipping {filename}: Invalid JSON format") except Exception as e: print(f"Error processing {filename}: {str(e)}") # Example usage (replace paths with your actual directories) if __name__ == "__main__": INPUT_DIR = "./your_tweet_json_folder" BACKUP_DIR = "./duplicate_backups" # Set to None to delete duplicates directly deduplicate_tweet_files(INPUT_DIR, BACKUP_DIR)
Key Customizations & Tips
- Safety First: Always use the
backup_dirparameter for your first run—this lets you review duplicates before permanently deleting anything - Normalization: Adjust the
normalized_textline to catch more duplicates (e.g., lowercase conversion, removing extra line breaks) - Memory Efficiency: Using SHA-256 hashes instead of storing full tweet text is critical if you have thousands of long tweets—this keeps memory usage low
- Error Handling: The script skips invalid JSON files and those missing the
TEXTfield, so you don't have to pre-clean your dataset
内容的提问来源于stack exchange,提问作者Bilal Butt

