You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何用库高效提取句子内容填充matrixA、matrixB?

Hey there! Great question—let’s break down the best tools for your task of splitting sentences with placeholders into your target matrices.

This is hands down the most efficient and straightforward choice for your use case. Regular expressions are designed exactly for pattern matching like this, so they’ll let you quickly extract the text preceding each {matrixA} or {matrixB} placeholder without writing tons of custom logic.

Here’s a working code example that does exactly what you need:

import re

matrixA = []
matrixB = []

sentences = [
    "words1 words2 words3 {matrixA} {matrixB}",
    "words3 words4 {matrixA}"
]

# Define a regex pattern to capture text before each placeholder
pattern = r"(.*?)\s*\{(matrix[A-B])\}"

for sent in sentences:
    # Find all matches of text + placeholder in the sentence
    matches = re.findall(pattern, sent)
    for text_segment, matrix_label in matches:
        # Clean up extra whitespace and add to the correct matrix
        cleaned_text = text_segment.strip()
        if matrix_label == "matrixA":
            matrixA.append(cleaned_text)
        elif matrix_label == "matrixB":
            matrixB.append(cleaned_text)

print(matrixA)  # Output: ["words1 words2 words3", "words3 words4"]
print(matrixB)  # Output: ["words1 words2 words3"]

Why this works:

  • The pattern (.*?)\s*\{(matrix[A-B])\} captures two groups:
    1. (.*?): The text before the placeholder (non-greedy match to stop at the first placeholder)
    2. (matrix[A-B]): The placeholder name (either matrixA or matrixB)
  • It handles sentences with multiple placeholders (like your first example) seamlessly, adding the same preceding text to both matrices if needed.

Honorable Mention: nltk (Less Ideal But Doable)

While nltk is a powerhouse for natural language processing tasks like tokenization and parsing, it’s overkill for this specific job. That said, if you’re already using nltk in your project, you can make it work:

import nltk
nltk.download('punkt')  # Only needed once

matrixA = []
matrixB = []

sentences = [
    "words1 words2 words3 {matrixA} {matrixB}",
    "words3 words4 {matrixA}"
]

for sent in sentences:
    # Split the sentence into tokens
    tokens = nltk.word_tokenize(sent)
    # Find indices of all placeholders
    placeholder_positions = [i for i, token in enumerate(tokens) if token.startswith('{') and token.endswith('}')]
    
    for idx in placeholder_positions:
        matrix_name = tokens[idx].strip('{}')
        # Join tokens before the placeholder into a single string
        preceding_text = ' '.join(tokens[:idx]).strip()
        
        if matrix_name == "matrixA":
            matrixA.append(preceding_text)
        elif matrix_name == "matrixB":
            matrixB.append(preceding_text)

print(matrixA)
print(matrixB)

Caveats with nltk:

  • You need to download the tokenization model first (punkt).
  • It’s less efficient than re for this specific pattern-matching task, since it’s doing full tokenization instead of targeted pattern matching.

Final Recommendation

Stick with re—it’s lightweight, fast, and perfectly tailored to your needs. Manual implementation is possible, but using re will save you time and reduce the chance of bugs in your custom parsing logic.

内容的提问来源于stack exchange,提问作者Budi Mulyo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:31:36