Linux下如何递归提取分层目录数据并导入MongoDB?
Got it, let's tackle this step by step. Dealing with hundreds of nested directories and mixed encodings can be tricky, but we can break it down into manageable parts:
1. 遍历所有文件并提取有效内容(处理编码问题)
First, we need to crawl through all those nested data/ subdirectories, read each file, and extract the email:password pairs—while handling both UTF-8 and Non-ISO extended-ASCII encodings. Python is perfect for this because it’s flexible with file I/O and encoding fallback.
Here’s a sample script that does this:
import os def extract_credentials(root_dir): credentials = [] # Walk through all directories and files recursively for dirpath, _, filenames in os.walk(root_dir): for filename in filenames: file_path = os.path.join(dirpath, filename) # Try UTF-8 first, fall back to latin-1 for extended ASCII files try: with open(file_path, 'r', encoding='utf-8') as f: lines = f.readlines() except UnicodeDecodeError: with open(file_path, 'r', encoding='latin-1') as f: lines = f.readlines() # Process each line to extract email:password pairs for line in lines: line = line.strip() if not line: continue # Skip empty lines # Split on the first colon (handles passwords with colons) parts = line.split(':', 1) if len(parts) == 2: email, password = parts credentials.append({ "email": email.strip(), "password": password.strip() }) return credentials # Run the extraction on your data directory all_credentials = extract_credentials('./data')
2. 导出为JSON或CSV
Once we have all credentials in a structured list, exporting to your desired format is straightforward.
导出为JSON
JSON is ideal for direct MongoDB import since it matches document structure:
import json with open('credentials.json', 'w', encoding='utf-8') as f: json.dump(all_credentials, f, indent=2)
导出为CSV
CSV works if you need a human-readable tabular format. We’ll use the csv module to handle special characters like commas:
import csv with open('credentials.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.DictWriter(f, fieldnames=["email", "password"]) writer.writeheader() # Write column headers writer.writerows(all_credentials)
3. 导入到MongoDB
You have two reliable options here: using the command-line mongoimport tool (fast for large files) or inserting directly via Python.
方法1:使用mongoimport(推荐)
导入JSON文件
mongoimport --db your_database_name --collection your_collection_name --file credentials.json --jsonArray
--jsonArraytells MongoDB the file contains a list of documents.
导入CSV文件
mongoimport --db your_database_name --collection your_collection_name --file credentials.csv --type csv --headerline
--headerlineuses the first CSV row as field names (matches our "email"/"password" headers).
方法2:直接用Python插入(适合实时处理)
If you want to skip intermediate files and insert straight into MongoDB:
from pymongo import MongoClient # Connect to your MongoDB instance (adjust connection string if using remote) client = MongoClient('mongodb://localhost:27017/') db = client['your_database_name'] collection = db['your_collection_name'] # Bulk insert all credentials result = collection.insert_many(all_credentials) print(f"Successfully inserted {len(result.inserted_ids)} documents")
关键注意事项
- Large Datasets: If you’re dealing with millions of entries, modify the script to process files in batches (write/insert chunks instead of loading everything into memory).
- Data Cleaning: Add regex checks to validate email formats (e.g.,
r'^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+$') if you need to filter invalid entries. - Encoding Edge Cases: If
latin-1fails for some files, trycp1252(a common Windows extended ASCII encoding).
内容的提问来源于stack exchange,提问作者pibace

