You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中移除字符串内EDIFACT消息的正则表达式实现需求

Remove EDIFACT Messages from String with Python Regex

Got it, let's break down how to safely remove EDIFACT messages from a string using Python regex without accidentally deleting other content. EDIFACT messages follow a strict structure we can exploit to target only them.

Key Observations About EDIFACT Structure

EDIFACT messages have consistent, identifiable markers:

  • They start with the UNB segment (interchange header)
  • They end with the UNZ segment (interchange trailer)
  • Each segment in between ends with a single quote ('), and segments may be separated by whitespace (spaces, newlines, tabs, etc.)

The Regex Pattern

We can build a regex that strictly matches this structure to avoid false positives:

UNB(?:[^']+'\s*)*UNZ[^']+'

Let's break down what each part does:

  • UNB: Matches the mandatory start of an EDIFACT interchange
  • (?:[^']+'\s*)*: A non-capturing group that matches any number of intermediate segments:
    • [^']+: Matches all characters in a segment (excluding the closing single quote)
    • ': Matches the segment's closing quote
    • \s*: Accounts for any whitespace (including newlines) between segments
    • *: Allows zero or more intermediate segments (valid EDIFACT will have at least some)
  • UNZ[^']+': Matches the mandatory closing UNZ segment, including its content up to the final quote

Python Implementation

Here's a reusable function to clean your text:

import re

def remove_edifact_content(input_text):
    # Define the EDIFACT-matching regex pattern
    edifact_regex = r'UNB(?:[^']+'\s*)*UNZ[^']+'
    # Replace all matched EDIFACT messages with an empty string
    cleaned_text = re.sub(edifact_regex, '', input_text)
    # Optional: Clean up leftover extra whitespace
    cleaned_text = re.sub(r'\s+', ' ', cleaned_text).strip()
    return cleaned_text

# Test with your sample content
sample_input = """Random text before EDIFACT UNB+AHBI:1+.? ' UNB+IATB:1+6XPPC:ZZ+LHPPC:ZZ+940101:0950+1' UNH+1+PAORES:93:1:IA' MSG+1:45' IFT+3+XYZCOMPANY AVAILABILITY' ERC+A7V:1:AMD' IFT+3+NO MORE FLIGHTS' ODI' TVL+240493:1000::1220+FRA+JFK+DL+400+C' PDI++C:3+Y::3+F::1' !ERC+21198:EC' APD+74C:0:::6++++++6X' TVL+240493:1740::2030+JFK+MIA+DL+081+C' PDI++C:4' APD+EM2:0:1630::6+++++++DA' UNT+13+1' UNZ+1+1' More random text after EDIFACT"""

print(remove_edifact_content(sample_input))
# Output: "Random text before EDIFACT More random text after EDIFACT"

Why This Won't Accidentally Delete Other Content

This regex is intentionally strict: it only targets content that starts with UNB, follows the segment structure (each part ending with '), and ends with a UNZ segment. Any other text—even if it contains single quotes or random letters—won't match unless it fits this exact EDIFACT pattern.

内容的提问来源于stack exchange,提问作者Deepak Aggarwal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:36:41