使用awk处理多文件:统计区间内点数的语法问题求助
Hey there! Let's walk through your AWK script issue step by step—this is a super common snag for new AWK users, so you’re not alone here.
The Problem Context
You’re trying to count how many points (from the points file) fall into each interval defined in the windows file. Your first script threw a syntax error near unexpected token 'i', but after tweaking it, you got the expected output. Let’s break down why the error happened and how the corrected code works.
Common Cause of the Initial Syntax Error
That error almost always comes from forgetting to wrap your entire AWK code in single quotes when running it in the shell. Without quotes, the shell tries to interpret AWK variables like i as shell variables, which breaks the script’s syntax. For example, writing:
awk NR==FNR{ name[NR]=$1 pos[NR]=$2 next } { for(i in name){ ... } } points windows
Instead of wrapping the code in single quotes causes the shell to misparse i, leading to the "unexpected token" error.
The Corrected Working Code
Here’s the properly formatted script that delivers the right results:
NR==FNR{ name[NR] = $1 pos[NR] = $2 next } { for(i in name){ if(name[i] == $1 && pos[i] >= $2 && pos[i] <= $3){ sum[FNR] += 1 } } } END { for(i = 1; i <= FNR; i++){ print sum[i]+0 # Ensures we print 0 for intervals with no matching points } }
Run it with this shell command (note the single quotes wrapping the AWK code):
awk 'NR==FNR{ name[NR]=$1; pos[NR]=$2; next } { for(i in name){ if(name[i]==$1 && pos[i]>=$2 && pos[i]<=$3){ sum[FNR]++ } } } END{ for(i=1; i<=FNR; i++){ print sum[i]+0 } }' points windows > output
How This Code Works
Let’s break down each section to make it clear:
Reading the
pointsfile:NR==FNRis a standard AWK trick to detect when we’re processing the first input file (points).NRis the total number of lines read so far, andFNRis the line number within the current file—they’re equal only for the first file.- We store each point’s chromosome name in the
namearray and its position in theposarray, using the line number as the index. nextskips the rest of the code, so this block only runs for lines inpoints.
Processing the
windowsfile:- For each interval (line in
windows), we loop through every point stored in our arrays. - We check two conditions: the point’s chromosome matches the interval’s chromosome, and the point’s position falls between the interval’s start (
$2) and end ($3). - If both conditions are true, we increment the counter
sum[FNR]—FNRhere is the line number of the current interval, so each interval gets its own unique counter.
- For each interval (line in
Output the results:
- The
ENDblock runs after all files are processed. We loop through thesumarray from 1 to the total number of intervals (which isFNRat the end of processingwindows). sum[i]+0ensures that even if an interval has no matching points (andsum[i]is uninitialized), we print0instead of nothing.
- The
Test Data & Expected Output
Your test files:
points
chr1 52 chr1 65 chr2 120 chr2 101 chr2 160 chr3 20 chr4 50
windows
chr1 0 100 chr1 100 200 chr2 0 100 chr2 100 200 chr3 0 100 chr3 100 200 chr4 0 100 chr5 0 100 chr6 0 100 chr6 100 200 chr7 0 100 chr8 0 100
The output file will contain:
2 0 0 3 1 0 1 0 0 0 0 0
This aligns with counting points per interval: 2 points in chr1 0-100, 3 in chr2 100-200, 1 in chr3 0-100, 1 in chr4 0-100, and 0 for all other intervals.
内容的提问来源于stack exchange,提问作者gpapprank

