为何基于多类别感知机的依存解析器反复选择同一动作?
问题描述
我正在从零开发一个基于多类别感知机的小型依存解析器,解析器包含START_ARC、MOVE_RIGHT、END_ARC三种动作用于处理待处理token。训练器可以通过输入学习到的动作准确复现树库中的依存弧,但ParserRunner.parse()方法中,感知机始终预测END_ARC动作,每个token都只会执行该动作,完全不选择其他动作。
依存解析器实现
class DependencyParser1: """ self.token_index must be manually increased along with the sentence, as each token has been processed. Do it as soon as self.start_arc() has been called. self.move_right() only moves the move pointer to the next token, while self.token_index keeps track of the token that is to be assigned a head. Each arc is a tuple [x, y] where x is the start, and y is the head. """ def __init__(self): self.move_index = 0 self.token_index = -1 self.arcs = [] self.arc = [0, 0] def move_right(self): self.move_index += 1 def start_arc(self): self.arc[0] = self.token_index + 1 def end_arc(self): self.arc[1] = 0 if self.move_index == 0 else self.move_index + 1 self.arcs.append(self.arc) self.arc = [0, 0] self.move_index = 0 def reset(self): self.arcs = [] self.arc = [0, 0] self.token_index = - 1 self.move_index = 0
训练器实现
class Trainer: def __init__(self, parser): self.parser = parser self.features = [] self.classes = [] def train(self, sentence): arc_start_index = 0 # 0 is the first token. The root # is numbered 0, the rest of the # tokens are numbered as the index # plus 1. sentence_length = len(sentence) while arc_start_index < sentence_length: self.parser.token_index += 1 self.parser.start_arc() self.features.append(FeaturesRecorder.record(locals())) self.classes.append(START_ARC) arc_end_index = 0 head = int(sentence[arc_start_index].head) - 1 while arc_end_index < head: arc_end_index += 1 self.parser.move_right() self.features.append(FeaturesRecorder.record(locals())) self.classes.append(MOVE_RIGHT) self.parser.end_arc() self.features.append(FeaturesRecorder.record(locals())) self.classes.append(END_ARC) arc_start_index += 1
解析运行器实现
class ParserRunner: def __init__(self, parser, perceptron): self.parser = parser self.perceptron = perceptron def cheat(self, sentence, actions): actions_copy = actions.copy() arc_start_index = 0 sentence_length = len(sentence) while actions_copy: action = actions_copy.pop(0) if action == START_ARC: self.parser.token_index += 1 self.parser.start_arc() elif action == MOVE_RIGHT: self.parser.move_right() elif action == END_ARC: self.parser.end_arc() return self.parser.arcs def parse(self, sentence): arc_start_index = 0 sentence_length = len(sentence) arc_is_started = 0 while arc_start_index < sentence_length: recording = FeaturesRecorder.record(locals()) print(recording) action = self.perceptron.predict([recording]) print(action) #print(action) if action == START_ARC: self.parser.token_index += 1 self.parser.start_arc() arc_is_started = 1 elif action == MOVE_RIGHT: self.parser.move_right() elif action == END_ARC: self.parser.end_arc() arc_start_index += 1 return self.parser.arcs
可能的排查方向
- 特征一致性问题:检查
FeaturesRecorder.record(locals())在训练和预测时提取的特征是否一致。训练时locals()包含arc_start_index、head、arc_end_index等变量,而预测时的locals()只有arc_start_index、arc_is_started等,缺少关键特征会导致感知机无法区分动作类别。 - 训练数据分布:统计训练集中三类动作的样本占比,如果
END_ARC样本占比过高,感知机会倾向于预测该类别。 - 感知机训练有效性:确认是否用
Trainer收集的features和classes完成了完整训练,权重更新逻辑是否正确执行。 - 解析逻辑不匹配:训练时每个token的处理流程是
START_ARC→若干MOVE_RIGHT→END_ARC,但预测时parse()方法中arc_start_index每次循环都会强制+1,不管当前执行的是什么动作。比如执行MOVE_RIGHT后,arc_start_index依然递增,导致后续token的处理状态和训练时完全不一致,特征错位。 - 状态重置问题:每次调用
parse()前,必须调用parser.reset()重置解析器状态,否则残留的状态会影响特征提取和动作预测。 - 动作编码一致性:确认
START_ARC、MOVE_RIGHT、END_ARC的编码(比如整数映射)在训练和预测时完全一致,编码不匹配会导致预测结果混乱。
内容的提问来源于stack exchange,提问作者Nils Blomqvist
相关产品推荐
相关产品推荐

