在AWS Ubuntu实例中批量解压百万+压缩包至指定目录结构
Hey there! As a fellow Linux user who's tackled bulk file processing headaches before, let's get this sorted for you. First, let's align on your directory structure—I’ll assume your setup looks something like this (adjust paths if yours differs):
./financial_data/Pair_USD_EUR/2024_01/transactions_001.zip ./financial_data/Pair_USD_EUR/2024_02/transactions_002.zip ./financial_data/Pair_GBP_JPY/2024_01/transactions_003.zip ...Each zip contains CSV files nested inside the month folder, and you want those CSVs extracted directly into the parent
Pair_XXXfolder, skipping the month subdirectory entirely.
Step 1: Test with a Single File First
Before running bulk commands, always test with one zip to confirm it works as expected:
# Replace this with a path to one of your zip files unzip -j ./financial_data/Pair_USD_EUR/2024_01/transactions_001.zip -d ./financial_data/Pair_USD_EUR/
-j: This "junks" the internal directory structure of the zip—so CSVs go straight to the target folder instead of recreating the month folder.-d: Specifies the destination directory (your Pair folder).
Check the Pair folder afterward—you should see the CSV files directly there, no extra month subfolder.
Step 2: Bulk Process All Zip Files
For 100k+ files, find is the most reliable tool to traverse your directories. Here's the command to process every zip and extract its CSVs to the parent Pair folder:
# Replace /path/to/your/root with the parent folder holding all your Pair directories find /path/to/your/root -name "*.zip" -exec sh -c ' zip_file="$1" # Get the parent Pair folder (go up two levels from the zip file) pair_dir=$(dirname "$(dirname "$zip_file")") # Extract CSVs directly to the Pair folder unzip -j "$zip_file" -d "$pair_dir" # Optional: Uncomment the line below to delete the zip after extraction (save space) # rm "$zip_file" ' sh {} \;
Key Notes:
- Adjust the path:
/path/to/your/rootmight be something like~/financial_dataor/mnt/ebs/financial_data(if you're using an EBS volume on AWS). - Duplicate files: If you have CSVs with the same name across zips,
unzipwill prompt you to overwrite. To auto-overwrite without prompts, add the-oflag:unzip -j -o "$zip_file" -d "$pair_dir".
Step 3: Speed Up Processing (Optional)
With 1 million files, single-threaded processing will take a while. Use xargs to run multiple extractions in parallel (match the number to your AWS instance's CPU cores—e.g., -P4 for 4 cores):
find /path/to/your/root -name "*.zip" | xargs -P4 -I {} sh -c ' zip_file="{}" pair_dir=$(dirname "$(dirname "$zip_file")") unzip -j "$zip_file" -d "$pair_dir" # rm "$zip_file" '
⚠️ Don't set -P too high—you'll risk overwhelming your instance's CPU or disk I/O. Start with a number equal to your core count and adjust if needed.
Critical Pre-Checks
- Disk Space: 1 million zips will unpack to a lot of data! Check your available space with
df -hbefore starting. - Permissions: Make sure you have read access to the zips and write access to the Pair folders. If not, run the commands with
sudo(but only if necessary).
内容的提问来源于stack exchange,提问作者swifty

