如何用KornShell和Awk遍历未知层级的子目录及文本文件?
Got it, let's work through this problem. You need a KornShell (ksh) script that traverses all levels of subdirectories starting from $FILEPATH, processes .txt files in a depth-first order (handle the deepest files first, then work back up to parent directories), and uses Awk for parsing/processing those files (for your database-related tasks). Here's a robust, easy-to-customize solution:
Core Script Implementation
This uses a recursive function to ensure we hit the deepest directories first before processing files in parent levels:
#!/usr/bin/ksh # Set your root directory here (can also pass as a script argument if needed) FILEPATH="/path/to/your/root/directory" # Recursive function to traverse directories in depth-first order traverse_depth_first() { local current_dir="$1" # First, process all subdirectories recursively (go deep first) for subdir in "$current_dir"/*/; do # Skip invalid directories (e.g., broken symlinks) if [ -d "$subdir" ]; then traverse_depth_first "$subdir" fi done # Now process all .txt files in the current directory (after subdirs are done) echo "=== Processing files in: $current_dir ===" for txt_file in "$current_dir"/*.txt; do # Only process actual files (avoids the "no files match" literal case) if [ -f "$txt_file" ]; then process_txt_with_awk "$txt_file" fi done } # Custom function to handle .txt files with Awk (tailor this to your needs) process_txt_with_awk() { local target_file="$1" echo " Working on: $target_file" # Example Awk logic (replace this with your database parsing commands) awk ' # Sample: Extract lines with database IDs and count them /^DB_ID:/ { ids[$2]++ print " Found ID:", $2 } END { print " Total unique IDs in", FILENAME, ":", length(ids) } ' "$target_file" } # Kick off the traversal from your root directory traverse_depth_first "$FILEPATH"
Key Details & Customization Tips
- Depth-First Logic: The recursive
traverse_depth_firstfunction first dives into every subdirectory, processing their contents (including deeper subdirs) before handling files in the current directory. This exactly matches your requirement of "deepest first, then back up". - Avoiding Empty Globs: The checks
[ -d "$subdir" ]and[ -f "$txt_file" ]prevent errors when a directory has no subdirs or no.txtfiles (without these, ksh would treat*.txtas a literal string if no files match). - Awk Integration: The
process_txt_with_awkfunction is where you'll plug in your actual database parsing logic. Modify the Awk script to extract data, filter records, generate outputs, or call other commands as needed. - Handling Hidden Directories: If you need to include directories starting with
., adjust the subdir glob to"$current_dir"/.*/and add a check to skip.and..(e.g.,if [ "$(basename "$subdir")" != "." ] && [ "$(basename "$subdir")" != ".." ]; then ...).
Alternative: Using ksh93's Globstar (If Depth Order Isn't Strict)
If your ksh version is ksh93 and you don't strictly need depth-first order (just need all .txt files regardless of level), you can use globstar for a simpler approach:
#!/usr/bin/ksh shopt -s globstar # Enable recursive globbing FILEPATH="/path/to/your/root/directory" for txt_file in "$FILEPATH"/**/*.txt; do [ -f "$txt_file" ] && process_txt_with_awk "$txt_file" done
Note: This traverses in breadth-first order (current directory first, then subdirs), so it won't match your "deepest first" requirement—but it's a handy shortcut if order doesn't matter.
内容的提问来源于stack exchange,提问作者Ausche

