You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Linux下如何递归提取分层目录数据并导入MongoDB?

处理分层目录数据集并导入MongoDB的完整方案

Got it, let's tackle this step by step. Dealing with hundreds of nested directories and mixed encodings can be tricky, but we can break it down into manageable parts:

1. 遍历所有文件并提取有效内容(处理编码问题)

First, we need to crawl through all those nested data/ subdirectories, read each file, and extract the email:password pairs—while handling both UTF-8 and Non-ISO extended-ASCII encodings. Python is perfect for this because it’s flexible with file I/O and encoding fallback.

Here’s a sample script that does this:

import os

def extract_credentials(root_dir):
    credentials = []
    # Walk through all directories and files recursively
    for dirpath, _, filenames in os.walk(root_dir):
        for filename in filenames:
            file_path = os.path.join(dirpath, filename)
            # Try UTF-8 first, fall back to latin-1 for extended ASCII files
            try:
                with open(file_path, 'r', encoding='utf-8') as f:
                    lines = f.readlines()
            except UnicodeDecodeError:
                with open(file_path, 'r', encoding='latin-1') as f:
                    lines = f.readlines()
            
            # Process each line to extract email:password pairs
            for line in lines:
                line = line.strip()
                if not line:
                    continue  # Skip empty lines
                # Split on the first colon (handles passwords with colons)
                parts = line.split(':', 1)
                if len(parts) == 2:
                    email, password = parts
                    credentials.append({
                        "email": email.strip(),
                        "password": password.strip()
                    })
    return credentials

# Run the extraction on your data directory
all_credentials = extract_credentials('./data')

2. 导出为JSON或CSV

Once we have all credentials in a structured list, exporting to your desired format is straightforward.

导出为JSON

JSON is ideal for direct MongoDB import since it matches document structure:

import json

with open('credentials.json', 'w', encoding='utf-8') as f:
    json.dump(all_credentials, f, indent=2)

导出为CSV

CSV works if you need a human-readable tabular format. We’ll use the csv module to handle special characters like commas:

import csv

with open('credentials.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.DictWriter(f, fieldnames=["email", "password"])
    writer.writeheader()  # Write column headers
    writer.writerows(all_credentials)

3. 导入到MongoDB

You have two reliable options here: using the command-line mongoimport tool (fast for large files) or inserting directly via Python.

方法1:使用mongoimport(推荐)

导入JSON文件

mongoimport --db your_database_name --collection your_collection_name --file credentials.json --jsonArray
  • --jsonArray tells MongoDB the file contains a list of documents.

导入CSV文件

mongoimport --db your_database_name --collection your_collection_name --file credentials.csv --type csv --headerline
  • --headerline uses the first CSV row as field names (matches our "email"/"password" headers).

方法2:直接用Python插入(适合实时处理)

If you want to skip intermediate files and insert straight into MongoDB:

from pymongo import MongoClient

# Connect to your MongoDB instance (adjust connection string if using remote)
client = MongoClient('mongodb://localhost:27017/')
db = client['your_database_name']
collection = db['your_collection_name']

# Bulk insert all credentials
result = collection.insert_many(all_credentials)
print(f"Successfully inserted {len(result.inserted_ids)} documents")

关键注意事项

  • Large Datasets: If you’re dealing with millions of entries, modify the script to process files in batches (write/insert chunks instead of loading everything into memory).
  • Data Cleaning: Add regex checks to validate email formats (e.g., r'^[a-zA-Z0-9_.+-]+@[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+$') if you need to filter invalid entries.
  • Encoding Edge Cases: If latin-1 fails for some files, try cp1252 (a common Windows extended ASCII encoding).

内容的提问来源于stack exchange,提问作者pibace

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:08:26