如何修改双向LSTM序列准备代码使输出偏移两个词?
修改BiLSTM输入输出序列以实现双单词偏移预测
你当前的代码是通过截断单词序列生成等长输入序列X,输出序列y对应输入偏移1个单词的结果。要实现偏移2个单词的预测,需要调整两个核心部分:
1. 适配双偏移的序列切片与边界补全
原来的代码仅补了1个<UNK>,但偏移2个单词时,需要补2个<UNK>避免索引越界;同时将y的切片起始位置调整为trunc_length + 2,确保每个y元素对应X对应位置往后第2个单词。
修改后的代码如下:
MAX_SEQ_LENGTH = 15 def prepare_sequences(words, unk_index, seqlen=MAX_SEQ_LENGTH): trunc_length = len(words) % seqlen # 生成输入序列:截断后分割为seqlen长度的块 X = np.array(words)[trunc_length:].reshape((-1, seqlen)) # 补2个<UNK>避免索引越界,起始位置偏移2个单词 y = np.array(words + [unk_index] * 2)[trunc_length + 2:].reshape((-1, seqlen)) return X, y Xtrain, ytrain = prepare_sequences(train_indices, word_to_index["<UNK>"]) Xtest, ytest = prepare_sequences(test_indices, word_to_index["<UNK>"]) # 检查形状,X与y的维度应保持一致 print(Xtrain.shape, ytrain.shape) print(Xtest.shape, ytest.shape)
2. 若需求为“输入一段序列,直接预测后续两个单词”
如果你的目标是用长度为15的单词序列,直接预测接下来的2个单词(每个输入样本对应2个输出标签),则需要调整序列分割逻辑,确保每个输入序列后有足够的单词作为输出:
MAX_SEQ_LENGTH = 15 PREDICT_NUM = 2 def prepare_sequences(words, unk_index, seqlen=MAX_SEQ_LENGTH, predict_num=PREDICT_NUM): # 计算截断长度,确保剩余单词可分割为(seqlen + predict_num)的完整块 total_segment_length = seqlen + predict_num trunc_length = len(words) % total_segment_length # 截断后分割为多个完整块 full_segments = np.array(words)[trunc_length:].reshape((-1, total_segment_length)) # X取每个块的前seqlen个单词,y取后predict_num个单词 X = full_segments[:, :seqlen] y = full_segments[:, seqlen:] # 处理最后一段不足的情况(可选补全) if len(words) - trunc_length < total_segment_length: padding_X = np.array(words[-seqlen:] + [unk_index]*(seqlen - len(words[-seqlen:]))) X = np.vstack([X, padding_X]) y = np.vstack([y, np.array([unk_index]*predict_num)]) return X, y Xtrain, ytrain = prepare_sequences(train_indices, word_to_index["<UNK>"]) Xtest, ytest = prepare_sequences(test_indices, word_to_index["<UNK>"]) # 此时X形状为(样本数, 15),y形状为(样本数, 2) print(Xtrain.shape, ytrain.shape) print(Xtest.shape, ytest.shape)
内容的提问来源于stack exchange,提问作者teri123
相关产品推荐
相关产品推荐

