SemEval 2017 Task-1语义文本相似度模型:连续目标值输入方法问询
Great question! Let’s break this down—you’re tackling a regression task (continuous similarity scores from 0 to 5) but using a 6-unit softmax output layer. This setup works by framing the regression problem as probabilistic regression, where the softmax outputs a probability distribution over the 0–5 integer anchors. Here’s exactly how to structure your target outputs to match this:
Core Idea
Softmax produces a normalized probability distribution across your 6 output units (each corresponding to the integer values 0, 1, 2, 3, 4, 5). So your target labels can’t be a single scalar—they need to be a 6-dimensional probability distribution that reflects where your continuous similarity score falls relative to those anchors.
Recommended Target Formats
1. Linear Interpolation (Soft Labels)
This is the simplest, most intuitive approach that preserves the continuous nature of your target scores:
- For any continuous target score
s(0 ≤ s ≤ 5):- Identify the two nearest integers:
floor_s = int(math.floor(s))andceil_s = int(math.ceil(s)) - Calculate weights based on how close
sis to each integer:weight_floor = 1 - (s - floor_s)weight_ceil = s - floor_s
- Build a 6-dimensional target vector where:
- The index matching
floor_sgetsweight_floor - The index matching
ceil_sgetsweight_ceil - All other indices get 0
- The index matching
- Identify the two nearest integers:
Example:
- If your target score is
2.3, the target vector is[0, 0, 0.7, 0.3, 0, 0] - If your target score is
5.0, the vector is[0, 0, 0, 0, 0, 1] - If your target score is
0.0, the vector is[1, 0, 0, 0, 0, 0]
2. Gaussian Kernel Smoothing (Smooth Probability Distributions)
For a more nuanced approach that captures uncertainty or gradual transitions between scores, use a Gaussian distribution to generate a smooth target probability distribution:
- For each output unit
i(corresponding to integeri, 0 ≤ i ≤5), calculate the target probability as:import math sigma = 0.5 # Adjust this to control distribution width (smaller = more focused) p_i = math.exp(-((s - i) ** 2) / (2 * sigma ** 2)) - Normalize all
p_ivalues so their sum equals 1 (to match the softmax output’s normalized property)
Example:
For s=2.3 and sigma=0.5, the probabilities for units 2 and 3 will be the highest, with smaller probabilities for adjacent units—creating a smooth bell curve centered at 2.3.
Pairing with Pearson Correlation Objective
Since you’re using Pearson correlation as your target function, you’ll need to convert the softmax output distribution back to a continuous predicted score first. Do this by computing the weighted sum of the 6 units:
predicted_score = 0*p0 + 1*p1 + 2*p2 + 3*p3 + 4*p4 + 5*p5
Where p0 to p5 are the softmax outputs for each unit. Then calculate the Pearson correlation between these predicted scores and your original continuous target scores, and optimize to maximize this correlation (or minimize 1 - correlation, since most optimizers are set up to minimize loss).
Key Notes
- Avoid hard one-hot labels (e.g., setting only unit 2 to 1 for a score of 2.3)—this discards valuable continuous information and will hurt model performance.
- In frameworks like PyTorch or TensorFlow, you can implement these target transformations with vectorized operations to keep things efficient.
内容的提问来源于stack exchange,提问作者shubham gupta

