You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何整理多行数据以适配scikit-learn RandomForest机器学习模型

问题:用机器学习模型定位RF发射器位置

我的目标是用机器学习模型确定射频(RF)发射器的位置,不想用三角测量、多接收器时间偏移这类传统技术,打算用scikit-learn的RandomForestClassifier来实现。

数据背景

  • 多个接收器位置固定,通过网络把数据传到中央数据库,接收器的RSSI值主要取决于和发射器之间是否有直视路径
  • RSSI范围是1-255,值为0表示未接收到信号,不会被记录到数据库;255是满量程信号
  • 数据库每秒记录一次数据,同一Time值代表某一时刻各接收器的信号情况:如果有X个接收器,同一时间点最少记录1行(只有1个接收器收到信号),最多记录X行(所有接收器都收到信号)
  • 已知位置数据是手动观察发射器在已知区域时的信号模式得到的,需要把这些信号模式和对应位置关联(比如Red、Green信号强对应Foo区域,Red、Yellow信号强但Blue弱对应Bar区域)

核心建模障碍

当前数据是按每个接收器的记录分行存储,同一时间点的信号分布在多行,且行数不固定,但RandomForestClassifier是逐行处理单一样本的。我需要按时间分组整理数据,但无法预知每个时间点有多少接收器能收到信号,不知道怎么把数据转换成适合建模的格式。

示例数据(Region A区域的几秒记录)

Receiver NameTimeRSSILocation
Red2024-03-21 20:37:58182Region A
Blue2024-03-21 20:37:58254Region A
Green2024-03-21 20:37:58208Region A
Red2024-03-21 20:37:59192Region A
Blue2024-03-21 20:37:59254Region A
Green2024-03-21 20:37:59215Region A
Red2024-03-21 20:38:00202Region A
Blue2024-03-21 20:38:00254Region A
Green2024-03-21 20:38:00207Region A
Yellow2024-03-21 20:38:0017Region A
Red2024-03-21 20:38:01189Region A
Blue2024-03-21 20:38:01254Region A
Green2024-03-21 20:38:01225Region A
Yellow2024-03-21 20:38:0116Region A
Red2024-03-21 20:38:02204Region A
Blue2024-03-21 20:38:02255Region A
Green2024-03-21 20:38:02213Region A
Yellow2024-03-21 20:38:0218Region A
Red2024-03-21 20:38:03180Region A
Blue2024-03-21 20:38:03254Region A
Green2024-03-21 20:38:03214Region A
Yellow2024-03-21 20:38:0313Region A
Red2024-03-21 20:38:04182Region A
Blue2024-03-21 20:38:04254Region A
Green2024-03-21 20:38:04213Region A
Yellow2024-03-21 20:38:0412Region A

现有代码(逐行处理,不确定是否正确)

我之前没用过Python和scikit-learn,写了这段初步代码,目前还是逐行处理数据:

import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder

data = pd.read_csv("data/combined.csv", header=0)

label_encoder = LabelEncoder()
print(data.columns)
data["name_encoded"] = label_encoder.fit_transform(data["name"])
data["location_encoded"] = label_encoder.fit_transform(data["location"])

x = data[["rssi", "name_encoded"]]  # Features (rssi and encoded name)
y = data["location_encoded"]  # Target (encoded location)

# Split data into training and testing sets
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42)

# Create the Random Forest model
model = RandomForestClassifier(n_estimators=100)

# Train the model
model.fit(x_train, y_train)

# Make predictions on the testing set
y_pred = model.predict(x_test)

# Decode predictions
location_decoder = LabelEncoder()

location_decoder.fit(data["location"])  # Fit the decoder with original locations
predicted_locations = location_decoder.inverse_transform(y_pred)
print("Predicted locations:", predicted_locations)

内容的提问来源于stack exchange,提问作者senfo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 21:49:51