You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Python中加载文本数据并构建2×n矩阵的实现方法

Create a 2×N Matrix from Q&A Text

Got it, let's work through how to build that 2-row matrix from your question-answer pair. The key goals are to split each sentence into individual tokens (words + punctuation), then make sure both rows have the same length by padding the shorter one with empty spaces.

Step-by-Step Breakdown

  1. Tokenize the Text: Split the question and answer into separate tokens, separating punctuation (like the ?) from words and converting everything to lowercase to match your example.
  2. Pad for Equal Length: Find the longest list of tokens, then add empty spaces to the shorter list until it matches that length.
  3. Build the Matrix: Combine the two padded lists into a 2-dimensional matrix.

Complete Python Solution

import re

def build_qa_matrix(question, answer):
    # Helper to split text into lowercase words and punctuation
    def tokenize(text):
        return re.findall(r'\w+|[^\w\s]', text.lower())
    
    # Get tokens for both Q and A
    q_tokens = tokenize(question)
    a_tokens = tokenize(answer)
    
    # Figure out how long each row needs to be
    max_row_length = max(len(q_tokens), len(a_tokens))
    
    # Pad the shorter row with spaces to match the longer one
    padded_q = q_tokens + [' '] * (max_row_length - len(q_tokens))
    padded_a = a_tokens + [' '] * (max_row_length - len(a_tokens))
    
    # Return the 2xN matrix
    return [padded_q, padded_a]

# Test with your input
question = "hello what is your name?"
answer = "Hi my name is John Smith"
matrix = build_qa_matrix(question, answer)
print(matrix)

Example Output

If you run this code, you'll get:

[['hello', 'what', 'is', 'your', 'name', '?', ' '], ['hi', 'my', 'name', 'is', 'john', 'smith', ' ']]

If you only want to pad the shorter row (instead of both when lengths are equal), you can tweak the padding logic to check which list is shorter first:

# Alternative padding logic (only pad the shorter row)
if len(q_tokens) < max_row_length:
    padded_q = q_tokens + [' '] * (max_row_length - len(q_tokens))
    padded_a = a_tokens
else:
    padded_a = a_tokens + [' '] * (max_row_length - len(a_tokens))
    padded_q = q_tokens

This way, if one row is naturally longer, only the shorter one gets the extra space(s) added.

内容的提问来源于stack exchange,提问作者Roberto Guglielmo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:24:24