如何使用Bash脚本比较本地文件夹文件并校验SQL生成文件的指定列名一致性
Got it, let's break down your two technical needs step by step—both are super practical when working with SQL-generated files saved locally!
The approach depends on what exactly you want to compare:
Quickly check for metadata differences (filenames, sizes, timestamps)
Use thediffcommand with recursive and quiet flags to spot discrepancies between two folders:diff -qr /path/to/folder1 /path/to/folder2-qhides detailed content differences and only tells you which files differ,-rchecks subfolders recursively.Compare file content for exact matches
For text files (like CSV/TSV), usediffdirectly:diff file1.csv file2.csvFor faster binary comparison (works for any file type), use
cmp:cmp file1.csv file2.csvIf your SQL exports might have rows in different order but identical content, sort first before comparing:
sort file1.csv > sorted1.csv && sort file2.csv > sorted2.csv && diff sorted1.csv sorted2.csvVisual side-by-side comparison
If you prefer a GUI tool,meldorkdiff3let you visually compare folders and files with color-coded differences.
Assuming your files are CSV/TSV with column names in the first row, here's a robust script to check that all files contain the required columns and have consistent column structures:
#!/bin/bash # Define required columns (adjust case sensitivity as needed) REQUIRED_COLS=("NAME" "COUNTRY" "STATE" "PHNO") # Replace with your target folder path TARGET_DIR="./sql_exported_files" # Check if target directory exists if [ ! -d "$TARGET_DIR" ]; then echo "Error: Directory $TARGET_DIR doesn't exist!" exit 1 fi # Get all CSV files (modify wildcard for other formats like *.tsv) FILES=("$TARGET_DIR"/*.csv) # Check if any files exist if [ ${#FILES[@]} -eq 0 ]; then echo "Error: No CSV files found in $TARGET_DIR!" exit 1 fi # Extract header from first file as the reference BASE_HEADER=$(head -n 1 "${FILES[0]}") # Split header into array (use IFS=$'\t' for TSV files) IFS=',' read -r -a BASE_COLS <<< "$BASE_HEADER" # First verify the reference file has all required columns echo "Validating required columns in reference file: ${FILES[0]}" for col in "${REQUIRED_COLS[@]}"; do if ! [[ " ${BASE_COLS[*]} " =~ " $col " ]]; then echo "Error: Reference file missing required column: $col" exit 1 fi done # Validate all other files against the reference echo "Checking all files for column consistency..." for file in "${FILES[@]}"; do echo "Processing: $file" CURRENT_HEADER=$(head -n 1 "$file") IFS=',' read -r -a CURRENT_COLS <<< "$CURRENT_HEADER" # Check if column count matches reference if [ ${#CURRENT_COLS[@]} -ne ${#BASE_COLS[@]} ]; then echo "⚠️ Warning: $file has a different number of columns than the reference!" fi # Check if columns match exactly (including order) for idx in "${!BASE_COLS[@]}"; do if [ "${CURRENT_COLS[$idx]}" != "${BASE_COLS[$idx]}" ]; then echo "❌ Error: $file column mismatch at position $idx: expected '${BASE_COLS[$idx]}', got '${CURRENT_COLS[$idx]}'" fi done # Optional: If column order doesn't matter, uncomment this to check required columns exist (ignore order) # for col in "${REQUIRED_COLS[@]}"; do # if ! [[ " ${CURRENT_COLS[*]} " =~ " $col " ]]; then # echo "❌ Error: $file missing required column: $col" # fi # done done echo "✅ All files passed column validation!"
Key Customizations for Your Use Case:
- Column Delimiter: If using TSV files, change
IFS=','toIFS=$'\t'. - Case Insensitivity: To ignore case differences, modify the comparison line to:
if [ "${CURRENT_COLS[$idx],,}" != "${BASE_COLS[$idx],,}" ]; then - File Format: Adjust the wildcard (
*.csv) to match your actual file extensions (e.g.,*.txt).
内容的提问来源于stack exchange,提问作者Gyari

