如何在Notepad++中查找文件内包含下划线的重复单词
Got it, let's figure out how to find those duplicate underscore-containing words across your entire file. From your example with #define directives, it sounds like you're targeting identifiers like FF_EIS_DMC_1_TERRAIN_AHEAD_TIME that show up more than once. Here are a few practical, tested methods to solve this:
1. Command Line: Grep + Sort + Uniq (Classic Approach)
This is a go-to for text processing tasks, especially if you're working in a terminal. It breaks down the problem into extracting, sorting, and filtering duplicates:
For your #define-specific scenario:
If your lines follow the exact format #define [number] [identifier] [comment], use this command to target only those underscore-containing identifiers:
grep -E '#define [0-9]+ ([A-Z_0-9]+)' your_file.h | awk '{print $3}' | sort | uniq -d
Let's break this down:
grep -Euses extended regex to match valid #define linesawk '{print $3}'pulls out the third field (your target identifier with underscores)sortgroups duplicate identifiers togetheruniq -donly outputs lines that appear more than once
For a general case (any underscore-containing word in the file):
If you need to scan the entire file for any word with underscores that repeats, use:
grep -oE '[A-Za-z0-9_]+' your_file | grep '_' | sort | uniq -d
grep -oEextracts individual words made of letters, numbers, and underscores- The second
grep '_'filters to only keep words with underscores sort+uniq -dhandles finding duplicates as before
2. Command Line: Awk (Single-Tool Efficiency)
Awk can handle the entire workflow in one command, which is faster for large files. It counts occurrences on the fly and outputs duplicates directly:
#define-specific version:
awk '/^#define/ {if ($3 ~ /_/) count[$3]++} END {for (word in count) if (count[word] > 1) print word}' your_file.h
- We check for lines starting with
#define, then verify the third field has an underscore - We increment a counter for each identifier
- At the end, we print any identifier that was counted more than once
General version:
awk '/_/ {count[$0]++} END {for (word in count) if (count[word] > 1) print word}' your_file
This scans every line, counts occurrences of any line containing an underscore, and outputs duplicates. Adjust the regex /[_]/ if you need to target full words instead of lines.
3. GUI Editor: VS Code (Visual Approach)
If you prefer a graphical tool, VS Code's regex-powered search can find duplicates across the entire file:
- Open your file and press
Ctrl+F(orCmd+Fon Mac) - Click the
.*button to enable regex mode - Paste this regex into the search bar:
\b([A-Za-z0-9_]+)\b(?=.*\b\1\b) - (Optional) Check "Match Case" if your identifiers are all uppercase, like in your example
- VS Code will highlight every instance of a duplicate underscore-containing word
This regex works by matching a word, then checking if that same word appears later in the file (using a positive lookahead and backreference \1).
内容的提问来源于stack exchange,提问作者Maverick283

