Python中移除字符串内EDIFACT消息的正则表达式实现需求
Got it, let's break down how to safely remove EDIFACT messages from a string using Python regex without accidentally deleting other content. EDIFACT messages follow a strict structure we can exploit to target only them.
Key Observations About EDIFACT Structure
EDIFACT messages have consistent, identifiable markers:
- They start with the
UNBsegment (interchange header) - They end with the
UNZsegment (interchange trailer) - Each segment in between ends with a single quote (
'), and segments may be separated by whitespace (spaces, newlines, tabs, etc.)
The Regex Pattern
We can build a regex that strictly matches this structure to avoid false positives:
UNB(?:[^']+'\s*)*UNZ[^']+'
Let's break down what each part does:
UNB: Matches the mandatory start of an EDIFACT interchange(?:[^']+'\s*)*: A non-capturing group that matches any number of intermediate segments:[^']+: Matches all characters in a segment (excluding the closing single quote)': Matches the segment's closing quote\s*: Accounts for any whitespace (including newlines) between segments*: Allows zero or more intermediate segments (valid EDIFACT will have at least some)
UNZ[^']+': Matches the mandatory closingUNZsegment, including its content up to the final quote
Python Implementation
Here's a reusable function to clean your text:
import re def remove_edifact_content(input_text): # Define the EDIFACT-matching regex pattern edifact_regex = r'UNB(?:[^']+'\s*)*UNZ[^']+' # Replace all matched EDIFACT messages with an empty string cleaned_text = re.sub(edifact_regex, '', input_text) # Optional: Clean up leftover extra whitespace cleaned_text = re.sub(r'\s+', ' ', cleaned_text).strip() return cleaned_text # Test with your sample content sample_input = """Random text before EDIFACT UNB+AHBI:1+.? ' UNB+IATB:1+6XPPC:ZZ+LHPPC:ZZ+940101:0950+1' UNH+1+PAORES:93:1:IA' MSG+1:45' IFT+3+XYZCOMPANY AVAILABILITY' ERC+A7V:1:AMD' IFT+3+NO MORE FLIGHTS' ODI' TVL+240493:1000::1220+FRA+JFK+DL+400+C' PDI++C:3+Y::3+F::1' !ERC+21198:EC' APD+74C:0:::6++++++6X' TVL+240493:1740::2030+JFK+MIA+DL+081+C' PDI++C:4' APD+EM2:0:1630::6+++++++DA' UNT+13+1' UNZ+1+1' More random text after EDIFACT""" print(remove_edifact_content(sample_input)) # Output: "Random text before EDIFACT More random text after EDIFACT"
Why This Won't Accidentally Delete Other Content
This regex is intentionally strict: it only targets content that starts with UNB, follows the segment structure (each part ending with '), and ends with a UNZ segment. Any other text—even if it contains single quotes or random letters—won't match unless it fits this exact EDIFACT pattern.
内容的提问来源于stack exchange,提问作者Deepak Aggarwal

