You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP中RubixML模型始终返回相同预测结果的问题求助

问题与解决方案:Rubix ML提取价格始终返回首个标签值

我尝试使用Rubix ML提取不同语句中的产品价格,但模型始终返回设置的第一个标签值260。已调整KNearestNeighbors的K值、扩充训练数据集、更换Vectorizer,但问题仍未解决。

原问题代码

<?php
include_once '../vendor/autoload.php';

use Rubix\ML\Datasets\Labeled;
use Rubix\ML\Datasets\Unlabeled;
use Rubix\ML\Classifiers\KNearestNeighbors;
use Rubix\ML\Transformers\WordCountVectorizer;
use Rubix\ML\Transformers\TfIdfTransformer;
use Rubix\ML\Pipeline;
use Rubix\ML\Extractors\CSV;

$samples= ['The price is 260','The cost is 500','This shirt costs 300','The value of this item is 450','Sold for 150 dolars'];
$labels = ['260',             '500',            '300',                 '450',                           '150'];

$dataset = new Labeled($samples, $labels);

// 生成模型
$pipeline = new Pipeline([
    new WordCountVectorizer(100),
    new TfIdfTransformer(),
], new KNearestNeighbors(3));

// 训练数据集
$pipeline->train($dataset);

// 分析新语句
$new = Unlabeled::build([
    ['Price: 1200'],
]);

// 预测
$predictions = $pipeline->predict($new);
var_dump($predictions);

问题根源

  1. 任务类型错配:你使用的KNearestNeighbors是分类模型,它会把每个价格当作独立类别,而非连续数值。模型核心逻辑是匹配文本相似度,而非识别提取数字。
  2. 特征局限性:当前的词频向量器仅关注词汇出现频率,无法将字符串中的数字作为关键特征。新输入的Price: 1200与训练集中The price is 260文本相似度最高,因此模型返回首个匹配的类别260。

可行解决方案

方案1:正则表达式直接提取(最高效)

价格提取属于有明确模式的任务,用正则比机器学习更直接可靠:

<?php
function extractPrice(string $text): ?int {
    // 匹配字符串中的整数价格
    preg_match('/\b\d+\b/', $text, $matches);
    return isset($matches[0]) ? (int)$matches[0] : null;
}

// 测试示例
echo extractPrice('Price: 1200'); // 输出 1200
echo extractPrice('The cost is 500'); // 输出 500
echo extractPrice('No price here'); // 输出 null

方案2:改用回归模型适配数值任务

如果必须使用机器学习,需将任务改为回归(预测连续数值),并调整模型类型:

<?php
include_once '../vendor/autoload.php';

use Rubix\ML\Datasets\Labeled;
use Rubix\ML\Datasets\Unlabeled;
use Rubix\ML\Regressors\KNearestNeighbors; // 注意使用回归版本的KNN
use Rubix\ML\Transformers\WordCountVectorizer;
use Rubix\ML\Transformers\TfIdfTransformer;
use Rubix\ML\Pipeline;

// 将标签改为数值类型(而非字符串)
$samples= ['The price is 260','The cost is 500','This shirt costs 300','The value of this item is 450','Sold for 150 dolars'];
$labels = [260, 500, 300, 450, 150];

$dataset = new Labeled($samples, $labels);

// 替换为回归版本的KNN模型
$pipeline = new Pipeline([
    new WordCountVectorizer(100),
    new TfIdfTransformer(),
], new KNearestNeighbors(3));

$pipeline->train($dataset);

$new = Unlabeled::build([['Price: 1200']]);

$predictions = $pipeline->predict($new);
var_dump($predictions); // 返回基于文本相似度的数值预测结果

注意:即使改用回归模型,文本特征对数值提取的效果依然有限。若要进一步提升精度,需额外加入数字特征提取逻辑,比如将字符串中的数字单独作为特征输入模型。

内容的提问来源于stack exchange,提问作者Pablo Arlia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 21:56:08