如何提取指定格式的document ID?支持document与ID间可变间隔
Got it, let's figure out how to extract the exact document ID you need from that log string. Here's a straightforward approach using regular expressions, tailored to your specific requirement:
We'll use a regex pattern that targets the sequence document → (any characters in between) → ID, then captures the first following UUID that matches your specified format.
Step 1: Regex Pattern
The pattern we'll use is:
document.*?ID.*?([a-fA-F0-9]{8}-[a-fA-F0-9]{4}-[a-fA-F0-9]{4}-[a-fA-F0-9]{4}-[a-fA-F0-9]{12})
Let's break down what each part does:
document.*?ID: Matches the literal string "document", followed by any number of characters (non-greedily) until it hits "ID" — this ensures we lock onto the correct section of the log..*?: Non-greedy match for any characters after "ID", so we don't skip past the first valid ID.- The parenthesized section
([a-fA-F0-9]{8}-...): Captures the UUID that fits your 8-4-4-4-12 alphanumeric format.
Step 2: Example Implementation (Python)
Here's how you can use this pattern in code to extract the ID:
import re # Your input log text log_content = "EventTimestamp H 9EventType 8document 2ID 2b837c02-40c9-4d33-b81b-d489a06fa302-DCUP LogToAuditTrail SourceAppCD 5DOCSV SourceAppUID 2b837c02-40c9-4d33-b81b-d489a06fa302 6UserID 5a8ce656-1a31-456b-b3dd-5ec0859c9f3e1" # Define the regex pattern doc_id_pattern = r'document.*?ID.*?([a-fA-F0-9]{8}-[a-fA-F0-9]{4}-[a-fA-F0-9]{4}-[a-fA-F0-9]{4}-[a-fA-F0-9]{12})' # Search for the pattern in the log match_result = re.search(doc_id_pattern, log_content) if match_result: extracted_doc_id = match_result.group(1) print(f"Extracted Document ID: {extracted_doc_id}") else: print("No valid document ID found matching the criteria.")
Step 3: Output
Running this code will output:
Extracted Document ID: 2b837c02-40c9-4d33-b81b-d489a06fa302
This approach ignores other UUIDs in the log (like the one after SourceAppUID or UserID) because it only looks for IDs that come right after the document...ID sequence.
内容的提问来源于stack exchange,提问作者Anurag Sharma

