Python中如何用库高效提取句子内容填充matrixA、matrixB?
Hey there! Great question—let’s break down the best tools for your task of splitting sentences with placeholders into your target matrices.
Recommended Library: re (Regular Expressions)
This is hands down the most efficient and straightforward choice for your use case. Regular expressions are designed exactly for pattern matching like this, so they’ll let you quickly extract the text preceding each {matrixA} or {matrixB} placeholder without writing tons of custom logic.
Here’s a working code example that does exactly what you need:
import re matrixA = [] matrixB = [] sentences = [ "words1 words2 words3 {matrixA} {matrixB}", "words3 words4 {matrixA}" ] # Define a regex pattern to capture text before each placeholder pattern = r"(.*?)\s*\{(matrix[A-B])\}" for sent in sentences: # Find all matches of text + placeholder in the sentence matches = re.findall(pattern, sent) for text_segment, matrix_label in matches: # Clean up extra whitespace and add to the correct matrix cleaned_text = text_segment.strip() if matrix_label == "matrixA": matrixA.append(cleaned_text) elif matrix_label == "matrixB": matrixB.append(cleaned_text) print(matrixA) # Output: ["words1 words2 words3", "words3 words4"] print(matrixB) # Output: ["words1 words2 words3"]
Why this works:
- The pattern
(.*?)\s*\{(matrix[A-B])\}captures two groups:(.*?): The text before the placeholder (non-greedy match to stop at the first placeholder)(matrix[A-B]): The placeholder name (eithermatrixAormatrixB)
- It handles sentences with multiple placeholders (like your first example) seamlessly, adding the same preceding text to both matrices if needed.
Honorable Mention: nltk (Less Ideal But Doable)
While nltk is a powerhouse for natural language processing tasks like tokenization and parsing, it’s overkill for this specific job. That said, if you’re already using nltk in your project, you can make it work:
import nltk nltk.download('punkt') # Only needed once matrixA = [] matrixB = [] sentences = [ "words1 words2 words3 {matrixA} {matrixB}", "words3 words4 {matrixA}" ] for sent in sentences: # Split the sentence into tokens tokens = nltk.word_tokenize(sent) # Find indices of all placeholders placeholder_positions = [i for i, token in enumerate(tokens) if token.startswith('{') and token.endswith('}')] for idx in placeholder_positions: matrix_name = tokens[idx].strip('{}') # Join tokens before the placeholder into a single string preceding_text = ' '.join(tokens[:idx]).strip() if matrix_name == "matrixA": matrixA.append(preceding_text) elif matrix_name == "matrixB": matrixB.append(preceding_text) print(matrixA) print(matrixB)
Caveats with nltk:
- You need to download the tokenization model first (
punkt). - It’s less efficient than
refor this specific pattern-matching task, since it’s doing full tokenization instead of targeted pattern matching.
Final Recommendation
Stick with re—it’s lightweight, fast, and perfectly tailored to your needs. Manual implementation is possible, but using re will save you time and reduce the chance of bugs in your custom parsing logic.
内容的提问来源于stack exchange,提问作者Budi Mulyo

