在Python中加载文本数据并构建2×n矩阵的实现方法
Create a 2×N Matrix from Q&A Text
Got it, let's work through how to build that 2-row matrix from your question-answer pair. The key goals are to split each sentence into individual tokens (words + punctuation), then make sure both rows have the same length by padding the shorter one with empty spaces.
Step-by-Step Breakdown
- Tokenize the Text: Split the question and answer into separate tokens, separating punctuation (like the
?) from words and converting everything to lowercase to match your example. - Pad for Equal Length: Find the longest list of tokens, then add empty spaces to the shorter list until it matches that length.
- Build the Matrix: Combine the two padded lists into a 2-dimensional matrix.
Complete Python Solution
import re def build_qa_matrix(question, answer): # Helper to split text into lowercase words and punctuation def tokenize(text): return re.findall(r'\w+|[^\w\s]', text.lower()) # Get tokens for both Q and A q_tokens = tokenize(question) a_tokens = tokenize(answer) # Figure out how long each row needs to be max_row_length = max(len(q_tokens), len(a_tokens)) # Pad the shorter row with spaces to match the longer one padded_q = q_tokens + [' '] * (max_row_length - len(q_tokens)) padded_a = a_tokens + [' '] * (max_row_length - len(a_tokens)) # Return the 2xN matrix return [padded_q, padded_a] # Test with your input question = "hello what is your name?" answer = "Hi my name is John Smith" matrix = build_qa_matrix(question, answer) print(matrix)
Example Output
If you run this code, you'll get:
[['hello', 'what', 'is', 'your', 'name', '?', ' '], ['hi', 'my', 'name', 'is', 'john', 'smith', ' ']]
If you only want to pad the shorter row (instead of both when lengths are equal), you can tweak the padding logic to check which list is shorter first:
# Alternative padding logic (only pad the shorter row) if len(q_tokens) < max_row_length: padded_q = q_tokens + [' '] * (max_row_length - len(q_tokens)) padded_a = a_tokens else: padded_a = a_tokens + [' '] * (max_row_length - len(a_tokens)) padded_q = q_tokens
This way, if one row is naturally longer, only the shorter one gets the extra space(s) added.
内容的提问来源于stack exchange,提问作者Roberto Guglielmo
相关产品推荐
相关产品推荐

