如何用VBScript/VBA统计[x.xx.xx]模式并提取对应内容
Got it, let's break down how to solve this problem. You need to count how many times the pattern [x.xx.xx] (like your examples [4.01.070] and [4.01.072]) shows up in a text variable, plus extract each matched pattern along with the text that comes right after it. Here's a practical Python solution using the built-in re module:
Step-by-Step Code Implementation
import re # Replace this with your actual text variable text_to_process = """ Lorem ipsum dolor sit amet [4.01.070] Consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam [4.01.072] Quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur. """ # Regex pattern tailored to match [x.xx.xx] (adjust digit lengths if needed) pattern = r'\[(\d\.\d{2}\.\d{3})\]' # Get all matches with their positions in the text matches = list(re.finditer(pattern, text_to_process)) # Calculate total number of occurrences total_occurrences = len(matches) # Extract each pattern and its subsequent content extracted_results = [] for index, match in enumerate(matches): # Grab the inner content of the pattern (without square brackets) matched_pattern = match.group(1) # Get text after the match: stop at next pattern or end of text if index < len(matches) - 1: next_match_position = matches[index + 1].start() follow_text = text_to_process[match.end():next_match_position].strip() else: follow_text = text_to_process[match.end():].strip() extracted_results.append({ "pattern": matched_pattern, "following_content": follow_text }) # Output the results in a readable format print(f"Total occurrences of the pattern: {total_occurrences}") print("\nDetailed breakdown:") for idx, result in enumerate(extracted_results, 1): print(f"\n{idx}. Matched pattern: *{result['pattern']}*") print(f" Text after pattern: {result['following_content']}")
What This Does
- Regex Pattern: The
r'\[(\d\.\d{2}\.\d{3})\]'targets your specific format:\[and\]escape square brackets (since they’re special in regex)(\d\.\d{2}\.\d{3})captures the core sequence: 1 digit → dot → 2 digits → dot → 3 digits. Adjust this (e.g., use\d+for variable-length digits) if your pattern varies (like[12.34.5678]).
- Extracting Follow-Up Text: Using
re.finditer()gives us each match’s position, so we can cleanly grab the text between matches (or from the last match to the end of the text). - Readable Output: The final print statements format the total count and each result with emphasized pattern text for clarity.
Sample Output
When you run the code with the example text, you’ll get:
Total occurrences of the pattern: 2 Detailed breakdown: 1. Matched pattern: *4.01.070* Text after pattern: Consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. 2. Matched pattern: *4.01.072* Text after pattern: Quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat.
内容的提问来源于stack exchange,提问作者Srihari
相关产品推荐
相关产品推荐

