PowerShell脚本如何使用两个自定义掩码筛选指定格式PDF文件?
Solution: Use Regex Matching for Custom Filename Masks
Since the -Filter parameter in Get-ChildItem only supports simple wildcard patterns (like *TEXT-TXT-*), we need to use regular expressions (regex) to match your specific, structured filename masks. Here's how to implement this in PowerShell:
Step-by-Step Explanation & Script
First, let's translate your two filename masks into regex patterns:
- Mask 1:
##-##-#### XXXX-XXX-XXXX-XX.PDF→ Digits for#, non-digits forX - Mask 2:
##-##-#### XXXX-XXX-XXX-XXXX-XX.PDF→ Same date prefix, longer non-digit segment
Full Script
# Set your target folder path $targetPath = "C:\Path\To\Your\PDF\Folder" # Define the regex pattern that matches both filename masks $filenameRegex = '^\d{2}-\d{2}-\d{4} (?:[^\d]{4}-[^\d]{3}-[^\d]{4}-[^\d]{2}|[^\d]{4}-[^\d]{3}-[^\d]{3}-[^\d]{4}-[^\d]{2})\.PDF$' # Get all PDF files, filter using the regex, then process each match Get-ChildItem -Path $targetPath -Filter "*.PDF" -Recurse -Force | Where-Object { $_.Name -match $filenameRegex } | ForEach-Object { # Replace this with your actual后续 steps (processing the PDF) Write-Host "Processing matching file: $($_.FullName)" # Example actions you could add: # - Extract text from the PDF # - Copy/move the file to another location # - Run a third-party tool on the file }
Regex Pattern Breakdown
Let's unpack the regex so you understand exactly how it works:
^: Anchors the match to the start of the filename (prevents partial matches, like a valid prefix in a longer filename)\d{2}-\d{2}-\d{4}: Matches the date segment (##-##-####) where\drepresents any digit, and{n}means exactlynoccurrences: Matches the space separating the date and non-digit segments(?:...|...): A non-capturing group that lets us match either of the two non-digit structures:- First option:
[^\d]{4}-[^\d]{3}-[^\d]{4}-[^\d]{2}→ MatchesXXXX-XXX-XXXX-XX([^\d]= any non-digit character) - Second option:
[^\d]{4}-[^\d]{3}-[^\d]{3}-[^\d]{4}-[^\d]{2}→ MatchesXXXX-XXX-XXX-XXXX-XX
- First option:
\.PDF$: Anchors the match to the end of the filename, ensuring it ends with.PDF(PowerShell's-matchis case-insensitive by default, so.pdfwill also match)
Key Tips
- Using
-Filter "*.PDF"first narrows down the files to only PDFs, making the regex check more efficient than scanning every file in the folder. - If you need strict case sensitivity for the
.PDFextension, replace-matchwith-cmatch(case-sensitive match). - If
Xshould only be letters (not symbols), replace[^\d]with[A-Za-z]in the regex.
内容的提问来源于stack exchange,提问作者Sunfile
相关产品推荐
相关产品推荐

