如何从1开始统计连续Y的出现次数直至后续值为N?(KDB+示例)
Solution for Counting Consecutive Ys (Resetting on N)
Got it, let's break down how to solve this problem where we need to count consecutive Y values starting from 1, and reset to 0 every time an N appears. Your example makes the requirement crystal clear:
q)tab:([]x:`N`N`Y`N`N`Y`Y`Y`N`N`Y`Y`Y;expected_output:0 0 1 0 0 1 2 3 0 0 1 2 3)
You mentioned trying the cut function, which is a solid approach. Let's expand on that, plus show a more streamlined method using kdb+'s built-in rle (run-length encoding) function, which is perfect for handling consecutive value sequences.
Method 1: Using rle (Most Efficient)
rle groups consecutive identical values into pairs of (value, run-length). We can use this to generate our count sequence directly:
First, get the run-length encoding of your
xcolumn:q)r:rle tab.x q)r (`N`Y`N`Y`N`Y;2 1 2 3 2 3) // (values, lengths of each consecutive run)Iterate over each run to generate the count sequence:
- For
Nruns: generate a list of 0s with the same length as the run - For
Yruns: generate a sequence starting at 1, up to the length of the run (using1+til length)
q)counts:raze {if[x=`N; enlist y#0; enlist 1+til y]} each flip r q)counts 0 0 1 0 0 1 2 3 0 0 1 2 3- For
Add this as a column to your table:
q)update expected_output:counts from tab x expected_output ------------------ N 0 N 0 Y 1 N 0 N 0 Y 1 Y 2 Y 3 N 0 N 0 Y 1 Y 2 Y 3
Method 2: Using cut (As You Initially Tried)
Your manual cut approach works, but we can automate the split indices instead of hardcoding them:
Generate the split indices using the run-length encoding's cumulative sums:
q)r:rle tab.x q)cuts:0, sums r[1] // Cumulative sum of run lengths gives split points q)cuts[:-1] // We exclude the final length to avoid an empty last segment 0 2 3 5 8 10Split the
xcolumn using these indices:q)segments:cuts[:-1] cut tab.x q)segments `N`N ,`Y `N`N `Y`Y`Y `N`N `Y`Y`YProcess each segment to generate counts:
q)counts:raze {if[all x=`N; count[x]#0; 1+til count x]} each segments q)counts 0 0 1 0 0 1 2 3 0 0 1 2 3Update your table with the result, same as Method 1.
Both methods will give you the exact expected output you need. The rle method is more concise and efficient, especially for large datasets, since it avoids manual index management.
内容的提问来源于stack exchange,提问作者DanDan4561

