Bash嵌套循环遍历目录执行命令:最后目录重复循环4次
Hey there! Let's break down your problem into two solvable parts: extracting the unique segment from your fastq filenames with sed, and fixing that annoying bug where the last directory gets processed 4 times.
1. Extracting the Unique Filename Segment
Your filenames follow the pattern name_S#_L001_R1_001.fastq.gz—the unique part is everything before _L001. Here's a sed command that snips off the redundant suffix cleanly:
# Example usage with a sample filename echo "sampleA_S3_L001_R1_001.fastq.gz" | sed -E 's/(_L001_R[12]_001\.fastq\.gz)//'
What this does:
-Eenables extended regular expressions for cleaner syntax.s/(_L001_R[12]_001\.fastq\.gz)//matches the fixed suffix (accounting for both R1 and R2 reads) and replaces it with nothing, leaving only your uniquename_S#segment.
You can embed this directly into your loop to capture the unique name for each file.
2. Fixing the Repeated Last Directory Loop
The root cause of your last directory repeating 4 times is almost certainly one of two issues:
- Variable name collision: Using the same variable (like
f) for both your outer directory loop and inner file loop, overwriting the outer loop's value. - Not returning to the parent directory: After
cding into a subdirectory, you don't navigate back, which can mess with subsequent loop iterations.
Here are two robust fixes:
Fix Option 1: Use Distinct Variables & Reset Working Directory
This approach saves your original directory, uses unique variable names, and ensures you return to the parent after processing each subdirectory:
# Save your starting directory to return to later original_dir=$(pwd) # Outer loop: iterate over all items in your target path, filter for directories for dir in /completepath/* do # Only process if the item is a directory if [ -d "$dir" ]; then # Enter the directory (skip if cd fails) cd "$dir" || continue # Inner loop: process each fastq.gz file for file in *.fastq.gz do # Extract the unique filename segment unique_name=$(echo "$file" | sed -E 's/(_L001_R[12]_001\.fastq\.gz)//') # Add your processing command here (e.g., alignment, renaming) echo "Processing $unique_name in directory: $(basename "$dir")" done # Return to your original starting directory cd "$original_dir" fi done
Fix Option 2: Use a Subshell to Isolate Directory Changes
A cleaner alternative is to wrap the directory-specific code in a subshell (()). This way, cd only affects the subshell, and your main shell stays in the original directory—no need to manually navigate back:
for dir in /completepath/* do if [ -d "$dir" ]; then # Subshell: all commands here run in a separate shell context ( cd "$dir" || exit # Exit subshell if cd fails for file in *.fastq.gz do unique_name=$(echo "$file" | sed -E 's/(_L001_R[12]_001\.fastq\.gz)//') echo "Processing $unique_name in directory: $(basename "$dir")" # Add your processing commands here done ) # Subshell exits, main shell is still in original directory fi done
Why this fixes the repetition:
By using unique variable names (dir for directories, file for files) we avoid overwriting the outer loop's iterator. The subshell method adds extra isolation, so directory changes never affect the main loop's state.
内容的提问来源于stack exchange,提问作者K-MO

