You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用正则提取量子化学输出中重复的NEVPT2多文本块?

Fixing NEVPT2 Block Extraction for Multi-Job Quantum Chemistry Output

Got it, the issue here is classic regex greediness—your original pattern uses the greedy * quantifier, which makes it grab all content from the first NEVPT2 start marker straight to the very last CASSCF TIMINGS line, instead of stopping at the end of each individual NEVPT2 block.

The Fix: Switch to Non-Greedy Matching

Simply modify the capture group to use the non-greedy *? quantifier instead of *. This tells the regex to match the minimum number of characters needed to reach the next occurrence of your end marker, which will isolate each NEVPT2 block perfectly.

Here's the updated regex:

NEVPT2_Section = r"(?:AILFT MATRIX ELEMENTS \(NEVPT2\)\n-+\n\n)([\s\S]*?)(?:\n\n--------------\nCASSCF TIMINGS)"

Why This Works

  • [\s\S]*? matches any character (including newlines) until the first instance of your end pattern (\n\n--------------\nCASSCF TIMINGS) is found.
  • This ensures each match corresponds to exactly one JOB's NEVPT2 results, rather than lumping all 31 blocks together.

Optional: Make It More Robust

If your output has minor inconsistencies in whitespace (extra newlines, trailing spaces), you can add \s* to handle variable spacing around the separators:

NEVPT2_Section = r"(?:AILFT MATRIX ELEMENTS \(NEVPT2\)\n-+\s*)([\s\S]*?)(?:\s*--------------\nCASSCF TIMINGS)"

Example Usage

Here's how to implement this to extract and save each NEVPT2 block individually:

import re

# Read your large output file
with open("quantum_chem_output.txt", "r") as infile:
    full_content = infile.read()

# Use the updated regex to find all NEVPT2 blocks
pattern = r"(?:AILFT MATRIX ELEMENTS \(NEVPT2\)\n-+\n\n)([\s\S]*?)(?:\n\n--------------\nCASSCF TIMINGS)"
for job_num, match in enumerate(re.finditer(pattern, full_content), start=1):
    nevpt2_data = match.group(1)
    # Save each block to a separate file
    with open(f"NEVPT2_results_job_{job_num}.txt", "w") as outfile:
        outfile.write(nevpt2_data)
    print(f"Successfully extracted NEVPT2 block for JOB {job_num}")

This will give you 31 separate files, each containing the NEVPT2 results for a single JOB—exactly what you need!

内容的提问来源于stack exchange,提问作者HCSthe2nd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:55:42