You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式匹配需求:提取文件中PCOMP起始至下一个CQUAD4前的内容

Solution to Capture PCOMP Sections Up to Next CQUAD4

Got it, let's break down how to use regex to extract all PCOMP sections that run up to (but don't include) the next CQUAD4 (or the end of the file, if there's no CQUAD4 after the last PCOMP).

The Regex Pattern

Here's the regex you'll need:

(?s)\bPCOMP\b.*?(?=\bCQUAD4\b|$)

Let's break down each part:

  • (?s): Enables "dotall" mode, so the . character matches newline characters too. This is crucial if your file has line breaks between entries.
  • \bPCOMP\b: Matches the exact word PCOMP (word boundaries \b prevent partial matches like PCOMP123).
  • .*?: Non-greedy match of any character (the ? ensures it stops at the first occurrence of our stop condition, not the last).
  • (?=\bCQUAD4\b|$): Positive lookahead that checks for either the next CQUAD4 (again, using word boundaries) or the end of the file ($). This ensures we don't include the CQUAD4 itself in our match.

Example Implementation (Python)

Here's how you'd use this regex in Python to process your file content:

import re

# Your file content (replace with reading from a file if needed)
file_content = """CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234 PCOMP 123 123 123 123 123 123 123 123 123 123 123 1231 23 CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234 PCOMP 123 123 123 123 123 123 123 123 123 123 123 1231 23 432 234 2342 34 CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234 CQUAD4 123123 123 234"""

# Extract all matches
matches = re.findall(r'(?s)\bPCOMP\b.*?(?=\bCQUAD4\b|$)', file_content)

# Print results
for i, match in enumerate(matches, 1):
    print(f"### Extracted PCOMP Section {i}")
    print(match.strip())
    print("---")

Expected Output

Running this code will give you two matches:

  1. PCOMP 123 123 123 123 123 123 123 123 123 123 123 1231 23
  2. PCOMP 123 123 123 123 123 123 123 123 123 123 123 1231 23 432 234 2342 34

Key Notes

  • Non-greedy matching is critical: Using .* instead of .*? would make the regex match from the first PCOMP all the way to the very last CQUAD4, which isn't what you want.
  • Word boundaries: If your file might have entries like PCOMPXYZ or CQUAD4_123, the \b ensures we only match the exact standalone words.
  • Handling files: If you're reading from a file, replace the file_content variable with open("your_file.txt").read().

内容的提问来源于stack exchange,提问作者user9415015

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:57:53