能否用Notepad++批量检索800+Word文档指定内容并导出Excel格式?
Yes, but it requires a few key steps since Notepad++ doesn’t natively process Word (.docx) files. The difficulty is moderate—you’ll need to handle batch conversion of Word docs to plain text, then use regular expressions (regex) to extract the required data, followed by light formatting to get Excel-ready output.
Step 1: Batch Convert Word Docs to Plain Text
Notepad++ works with text files, so first you need to convert all 800+ .docx files to .txt. Here’s an efficient way using Microsoft Word:
- Open Word, enable the Developer tab (if hidden: go to File > Options > Customize Ribbon and check "Developer").
- Click Macros, name it
BatchConvertToTxt, then click Create. - Replace the default macro code with this:
Sub BatchConvertToTxt() Dim fd As FileDialog Dim inputPath As String Dim outputPath As String Dim doc As Document Set fd = Application.FileDialog(msoFileDialogFolderPicker) If fd.Show = -1 Then inputPath = fd.SelectedItems(1) & "\" outputPath = inputPath & "TextFiles\" MkDir outputPath inputFile = Dir(inputPath & "*.docx") Do While inputFile <> "" Set doc = Documents.Open(inputPath & inputFile) doc.SaveAs2 Filename:=outputPath & Replace(inputFile, ".docx", ".txt"), FileFormat:=wdFormatText doc.Close SaveChanges:=False inputFile = Dir() Loop End If End Sub
- Run the macro, select the folder with your Word docs, and it’ll create a
TextFilessubfolder with all converted .txt files.
Step 2: Extract Data with Notepad++’s "Find in Files"
Open Notepad++ and follow these steps:
- Go to Search > Find in Files.
- In the Find what field, enter this regex pattern (adjust if your docs have minor formatting variations):
(?s)^(\d+).*?^.*Map Zones\s*([\d,\s]+)
- Breakdown:
(?s): Lets.match newlines (single-line mode)^(\d+): Captures the first number at the start of the document’s first line.*?: Skips content non-greedily until the line with "Map Zones"^.*Map Zones\s*: Finds the line containing "Map Zones" and skips leading text/whitespace after the term([\d,\s]+): Captures the numbers, commas, and spaces following "Map Zones"
- Set Filters to
*.txt, Directory to yourTextFilesfolder. - Check Regular expression under Search Mode.
- Click Find All to generate a list of matches.
Step 3: Format Output for Excel
Clean up the results to match your desired format:
- Copy the Find in Files results to a new Notepad++ tab.
- Go to Search > Replace:
- Find what:
.*?(\d+).*?([\d,\s]+) - Replace with:
BpS Code Map Zones $1 $2 - Check Regular expression and click Replace All.
- Find what:
- Copy the cleaned lines into Excel, then use Data > Text to Columns with spaces as delimiters to split the content into separate columns.
Alternative: More Reliable with a Python Script
If your Word docs have inconsistent formatting, a Python script will handle it more smoothly (no conversion needed):
import os from docx import Document input_folder = "path/to/your/word/docs" output_file = "results.txt" with open(output_file, "w") as f: f.write("BpS Code\tMap Zones\n") # Tab-separated for easy Excel import for filename in os.listdir(input_folder): if filename.endswith(".docx"): doc = Document(os.path.join(input_folder, filename)) # Extract first number from the first paragraph first_paragraph = doc.paragraphs[0].text.strip() bps_code = next((word for word in first_paragraph.split() if word.isdigit()), "") # Extract Map Zones numbers map_zones = "" for para in doc.paragraphs: text = para.text.strip() if "Map Zones" in text: parts = text.split("Map Zones")[-1].strip() map_zones = ''.join(c for c in parts if c.isdigit() or c in (',', ' ')).strip() break if bps_code and map_zones: f.write(f"{bps_code}\t{map_zones}\n")
This creates a tab-separated file you can open directly in Excel. Install the required library first with pip install python-docx.
内容的提问来源于stack exchange,提问作者Randy Swaty

