基于文件名日期删除30天前文件的技术求助(CSV格式文件名)
Got it, let's work through this problem. Since you want to base the deletion on the date in the filename (not the file's modification time, which is what find -mtime +30 uses), we can build a custom script to parse those dates and set up a recurring task to run it every 30 days.
Step 1: Build a Script to Delete Files Based on Filename Date
First, we'll create a bash script that:
- Calculates the date 30 days prior to today
- Scans your CSV directory
- Extracts the date from each filename
- Compares it to the threshold and deletes old files
Example Script (for prefix_YYYYMMDD.csv format)
Let's assume your files look like sales_report_20240101.csv (date in YYYYMMDD format at the end, after an underscore). Adjust the date extraction logic to match your actual filename pattern.
#!/bin/bash # Replace this with the path to your CSV files TARGET_DIR="/home/youruser/csv_files" # Calculate the threshold date (30 days ago) in YYYYMMDD format THRESHOLD=$(date -d "30 days ago" +%Y%m%d) # Loop through all CSV files in the target directory for file in "$TARGET_DIR"/*.csv; do # Skip if the path isn't a valid file (e.g., no CSV files found) [ -f "$file" ] || continue # Get just the filename (without the full path) filename=$(basename "$file") # Extract the date from the filename (adjust this to match your pattern!) # For "sales_report_20240101.csv": remove everything before the last underscore, then remove .csv file_date=${filename#*_} file_date=${file_date%.csv} # Validate the extracted date is a valid 8-digit number (optional but safe) if ! [[ "$file_date" =~ ^[0-9]{8}$ ]]; then echo "Skipping $filename: invalid date format detected" continue fi # Compare dates: if the file's date is older than the threshold, delete it if [ "$file_date" -lt "$THRESHOLD" ]; then echo "Deleting old file: $file" # Uncomment the line below once you've tested and confirmed it works # rm "$file" fi done
Adjusting for Other Date Formats
If your filenames use YYYY-MM-DD (like report-2024-01-01.csv), modify these parts:
- Update the threshold to match the format:
THRESHOLD=$(date -d "30 days ago" +%Y-%m-%d) - Extract the date with a regex to catch the hyphenated format:
file_date=$(echo "$filename" | grep -oE '[0-9]{4}-[0-9]{2}-[0-9]{2}') - Compare using string comparison (since hyphenated dates are lex order-matching chronological order):
if [[ "$file_date" < "$THRESHOLD" ]]; then
Step 2: Set Up a Recurring 30-Day Task
Cron doesn't have a direct "every 30 days" option, but we have two reliable ways to set this up:
Option 1: Approximate Every 30 Days (Monthly Twice)
Run the script on the 1st and 21st of each month (roughly 30 days apart):
- Open your crontab editor:
crontab -e - Add this line (replace the script path with your actual file location):
This runs at midnight (00:00) on the 1st and 21st, and logs output to a file for debugging.0 0 1,21 * * /path/to/your/delete_old_csvs.sh >> /var/log/old_csv_deletions.log 2>&1
Option 2: Strict 30-Day Intervals
If you need exact 30-day gaps, use a cron job that checks the elapsed time since the last run. Alternatively, use anacron (if your system supports it) which handles intervals better for non-24/7 systems.
Key Notes
- Test First: Run the script manually first with the
rmline commented out to verify it lists the correct files to delete. - Permissions: Make the script executable with
chmod +x /path/to/your/delete_old_csvs.sh. - Special Characters: If your filenames have spaces or special characters, the script uses double quotes around variables to avoid breaking paths.
- Customize Extraction: Tweak the
file_dateextraction logic to match your specific filename pattern (e.g., if the date is at the start like20240101_sales.csv, usefile_date=${filename%_*}instead).
内容的提问来源于stack exchange,提问作者user9495954

