PHP中RubixML模型始终返回相同预测结果的问题求助
问题与解决方案:Rubix ML提取价格始终返回首个标签值
我尝试使用Rubix ML提取不同语句中的产品价格,但模型始终返回设置的第一个标签值260。已调整KNearestNeighbors的K值、扩充训练数据集、更换Vectorizer,但问题仍未解决。
原问题代码
<?php include_once '../vendor/autoload.php'; use Rubix\ML\Datasets\Labeled; use Rubix\ML\Datasets\Unlabeled; use Rubix\ML\Classifiers\KNearestNeighbors; use Rubix\ML\Transformers\WordCountVectorizer; use Rubix\ML\Transformers\TfIdfTransformer; use Rubix\ML\Pipeline; use Rubix\ML\Extractors\CSV; $samples= ['The price is 260','The cost is 500','This shirt costs 300','The value of this item is 450','Sold for 150 dolars']; $labels = ['260', '500', '300', '450', '150']; $dataset = new Labeled($samples, $labels); // 生成模型 $pipeline = new Pipeline([ new WordCountVectorizer(100), new TfIdfTransformer(), ], new KNearestNeighbors(3)); // 训练数据集 $pipeline->train($dataset); // 分析新语句 $new = Unlabeled::build([ ['Price: 1200'], ]); // 预测 $predictions = $pipeline->predict($new); var_dump($predictions);
问题根源
- 任务类型错配:你使用的
KNearestNeighbors是分类模型,它会把每个价格当作独立类别,而非连续数值。模型核心逻辑是匹配文本相似度,而非识别提取数字。 - 特征局限性:当前的词频向量器仅关注词汇出现频率,无法将字符串中的数字作为关键特征。新输入的
Price: 1200与训练集中The price is 260文本相似度最高,因此模型返回首个匹配的类别260。
可行解决方案
方案1:正则表达式直接提取(最高效)
价格提取属于有明确模式的任务,用正则比机器学习更直接可靠:
<?php function extractPrice(string $text): ?int { // 匹配字符串中的整数价格 preg_match('/\b\d+\b/', $text, $matches); return isset($matches[0]) ? (int)$matches[0] : null; } // 测试示例 echo extractPrice('Price: 1200'); // 输出 1200 echo extractPrice('The cost is 500'); // 输出 500 echo extractPrice('No price here'); // 输出 null
方案2:改用回归模型适配数值任务
如果必须使用机器学习,需将任务改为回归(预测连续数值),并调整模型类型:
<?php include_once '../vendor/autoload.php'; use Rubix\ML\Datasets\Labeled; use Rubix\ML\Datasets\Unlabeled; use Rubix\ML\Regressors\KNearestNeighbors; // 注意使用回归版本的KNN use Rubix\ML\Transformers\WordCountVectorizer; use Rubix\ML\Transformers\TfIdfTransformer; use Rubix\ML\Pipeline; // 将标签改为数值类型(而非字符串) $samples= ['The price is 260','The cost is 500','This shirt costs 300','The value of this item is 450','Sold for 150 dolars']; $labels = [260, 500, 300, 450, 150]; $dataset = new Labeled($samples, $labels); // 替换为回归版本的KNN模型 $pipeline = new Pipeline([ new WordCountVectorizer(100), new TfIdfTransformer(), ], new KNearestNeighbors(3)); $pipeline->train($dataset); $new = Unlabeled::build([['Price: 1200']]); $predictions = $pipeline->predict($new); var_dump($predictions); // 返回基于文本相似度的数值预测结果
注意:即使改用回归模型,文本特征对数值提取的效果依然有限。若要进一步提升精度,需额外加入数字特征提取逻辑,比如将字符串中的数字单独作为特征输入模型。
内容的提问来源于stack exchange,提问作者Pablo Arlia
相关产品推荐
相关产品推荐

