通道竞态条件排查:inotify事件过滤测试偶发死锁问题
Hey there, let's dig into this flaky deadlock issue you're hitting in your tests—random failures are the worst! Based on your setup with rfsnotify and channel-based event filtering, here are the most likely culprits and actionable fixes:
Chances are your eventFilter is running in a goroutine, but your test isn't properly syncing with its execution or ensuring the output channel gets closed. If the test tries to read from the channel before the goroutine has pushed the event (or if the goroutine gets stuck and never sends), you'll hit that occasional deadlock.
Try modifying your function to accept a done signal channel, and ensure the output channel is closed when the goroutine exits:
// Updated eventFilter with exit control func eventFilter(watcher *rfsnotify.RWatcher, excludes []string, out chan<- fsnotify.Event, done <-chan struct{}) { defer close(out) // Guarantee channel closure to prevent permanent read blocking for { select { case event, ok := <-watcher.Events: if !ok { return } // Run your exclusion logic here if !isExcluded(event.Path, excludes) { out <- event } case _, ok := <-watcher.Errors: if !ok { return } case <-done: // Exit gracefully when test signals completion return } } } // In your test: func TestEventFilter(t *testing.T) { watcher, err := rfsnotify.NewWatcher() require.NoError(t, err) defer watcher.Close() out := make(chan fsnotify.Event) done := make(chan struct{}) defer close(done) // Trigger exit signal when test finishes // Spin up the filter goroutine go eventFilter(watcher, []string{"ignore-me.txt"}, out, done) // Trigger an inotify event (e.g., create a test file) tempFile, err := os.CreateTemp("", "test-filter") require.NoError(t, err) defer os.Remove(tempFile.Name()) // Use a timed select to avoid permanent deadlock select { case event := <-out: assert.Equal(t, fsnotify.Create, event.Op) assert.Contains(t, event.Path, "test-filter") case <-time.After(500 * time.Millisecond): t.Fatal("Timed out waiting for filtered event—goroutine may be stuck or event was dropped") } }
Even if production runs smoothly, test environments can expose edge cases. Double-check your exclusion logic for any operations that might occasionally block (e.g., synchronous file stat calls, slow string matching on large paths). A blocked filter goroutine can't push events to the channel, leading to test timeouts.
Quick fix for testing: Use a minimal exclusion list (or empty list) to rule out filter logic as the culprit. If the deadlock stops, refactor the filter to avoid blocking operations.
Inotify events don't fire instantaneously—system load or kernel scheduling delays can cause tiny gaps between triggering an event and it reaching your watcher. If your test tries to read the channel immediately after triggering the event, it might beat the goroutine to the punch.
Adding a buffered channel can help here:
// Add a small buffer to avoid immediate blocking on event send out := make(chan fsnotify.Event, 1)
This lets the goroutine push the event without waiting for the test to start reading, even if there's a tiny delay.
If you want stricter control over goroutine execution, use a sync.WaitGroup to ensure the filter has finished processing events before your test exits:
func TestEventFilter(t *testing.T) { watcher, err := rfsnotify.NewWatcher() require.NoError(t, err) defer watcher.Close() out := make(chan fsnotify.Event, 1) done := make(chan struct{}) var wg sync.WaitGroup wg.Add(1) go func() { defer wg.Done() eventFilter(watcher, []string{}, out, done) }() // Trigger event... // Wait for event and validate select { case event := <-out: assert.True(t, event.Has(fsnotify.Create)) case <-time.After(1 * time.Second): t.Fatal("No event received") } // Signal filter to exit and wait for it to clean up close(done) wg.Wait() }
The core issue here is almost always missing synchronization between your test and the background goroutine handling events. Adding exit signals, timed reads, or explicit waits should eliminate that 30% failure rate.
内容的提问来源于stack exchange,提问作者rumdrums

