Shell脚本嵌套循环执行报错:EOF错误排查求助
Let's break down the issues in your script and fix them one by one, starting with the EOF error you mentioned.
1. Root Cause of the EOF Error
Your script is missing the closing done statement for the outer for f loop. When the shell reaches the end of the file without finding this closing tag, it throws an unexpected EOF error. That's the immediate fix you need to apply.
2. Other Critical Issues to Fix
Beyond the EOF error, there are several other problems that will cause your script to fail or behave unpredictably:
a. Avoid Using ls in For Loops
Using $(ls *.txt.sumstats.gz) can break if filenames contain spaces or special characters. Instead, use glob patterns directly in the loop to ensure reliability:
for f in *.txt.sumstats.gz; do
b. Uninitialized k Variable
The parallel counter k isn't initialized to 0 before your loops run. This will cause arithmetic errors the first time ((k=k%ncors)) executes. Add k=0 before the nested loops to fix this.
c. Repeated File Extraction
Your script unzips the same .gz files multiple times (once per inner loop iteration). After the first unzip, the .gz file no longer exists, so subsequent iterations will throw "file not found" errors. Extract all files once before the loops instead.
d. Incorrect Sed Pattern Syntax
The sed commands use \m and \M which aren't valid in sed's default basic regular expression mode. Remove the backslashes to match the literal m and M in your filenames:
fff=$(echo ${ff} | sed 's/munged_Summary_statistics_MAGNETIC_//') iii=$(echo ${ii} | sed 's/Munged_ENIGMA2_//')
e. Race Conditions with Background Processes
Running multiple background processes that modify the same files (like unzipping) can lead to race conditions. Handling file extraction outside the parallelized section avoids this risk entirely.
3. Fixed Script
Here's the revised script with all these fixes applied:
#!/bin/bash #SBATCH -N 1 #SBATCH -t 00:05:00 # Bash script run for LDSC dataset # Load modules module load 2019 module load Python/2.7.15-foss-2019b # Copy files to scratch cp -r $HOME/ldsc/eur_w_ld_chr $TMPDIR cp $HOME/output_dir/*_txt.sumstats.gz $TMPDIR cp $HOME/output_dir/*_Jan23.tbl.sumstats.gz $TMPDIR myLDSC=$HOME/ldsc/ldsc.py cd $TMPDIR # Create output dir mkdir -p "$TMPDIR"/output_dir ncors=15 k=0 # Initialize parallel counter # Extract all sumstats files once to avoid repeated extraction for f in *.txt.sumstats.gz; do gunzip "$f" done for i in *_Jan23.tbl.sumstats.gz; do gunzip "$i" done # Run LDSC in parallel with nested loops for f in *.txt.sumstats; do for i in *_Jan23.tbl.sumstats; do ((k=k%ncors)); ((k++==0)) && wait { # Extract base names without prefixes fff=$(echo "${f}" | sed 's/munged_Summary_statistics_MAGNETIC_//') iii=$(echo "${i}" | sed 's/Munged_ENIGMA2_//') # Run LDSC command "$myLDSC" --rg "${f},${i}" --out output_dir/LDSC_"${fff}"_"${iii}" --ref-ld-chr eur_w_ld_chr/ --w-ld-chr eur_w_ld_chr/ } & done done # This closing tag was missing, causing the EOF error wait # Copy output dir from scratch to home cp -r "$TMPDIR"/output_dir $HOME
4. Additional Notes
- If you need to preserve the original
.gzfiles, extract them to a temporary subdirectory instead of the main$TMPDIR, or check if LDSC supports reading compressed files directly (usingzcator built-in compression support). - Always quote variables (like
"$f"or"$myLDSC") to handle spaces or special characters in filenames/paths. - The final
waitensures all background processes finish before copying the output back to your home directory, preventing incomplete files from being copied.
内容的提问来源于stack exchange,提问作者Tim Van Der Es

