如何整理多行数据以适配scikit-learn RandomForest机器学习模型
问题:用机器学习模型定位RF发射器位置
我的目标是用机器学习模型确定射频(RF)发射器的位置,不想用三角测量、多接收器时间偏移这类传统技术,打算用scikit-learn的RandomForestClassifier来实现。
数据背景
- 多个接收器位置固定,通过网络把数据传到中央数据库,接收器的RSSI值主要取决于和发射器之间是否有直视路径
- RSSI范围是1-255,值为0表示未接收到信号,不会被记录到数据库;255是满量程信号
- 数据库每秒记录一次数据,同一
Time值代表某一时刻各接收器的信号情况:如果有X个接收器,同一时间点最少记录1行(只有1个接收器收到信号),最多记录X行(所有接收器都收到信号) - 已知位置数据是手动观察发射器在已知区域时的信号模式得到的,需要把这些信号模式和对应位置关联(比如Red、Green信号强对应Foo区域,Red、Yellow信号强但Blue弱对应Bar区域)
核心建模障碍
当前数据是按每个接收器的记录分行存储,同一时间点的信号分布在多行,且行数不固定,但RandomForestClassifier是逐行处理单一样本的。我需要按时间分组整理数据,但无法预知每个时间点有多少接收器能收到信号,不知道怎么把数据转换成适合建模的格式。
示例数据(Region A区域的几秒记录)
| Receiver Name | Time | RSSI | Location |
|---|---|---|---|
| Red | 2024-03-21 20:37:58 | 182 | Region A |
| Blue | 2024-03-21 20:37:58 | 254 | Region A |
| Green | 2024-03-21 20:37:58 | 208 | Region A |
| Red | 2024-03-21 20:37:59 | 192 | Region A |
| Blue | 2024-03-21 20:37:59 | 254 | Region A |
| Green | 2024-03-21 20:37:59 | 215 | Region A |
| Red | 2024-03-21 20:38:00 | 202 | Region A |
| Blue | 2024-03-21 20:38:00 | 254 | Region A |
| Green | 2024-03-21 20:38:00 | 207 | Region A |
| Yellow | 2024-03-21 20:38:00 | 17 | Region A |
| Red | 2024-03-21 20:38:01 | 189 | Region A |
| Blue | 2024-03-21 20:38:01 | 254 | Region A |
| Green | 2024-03-21 20:38:01 | 225 | Region A |
| Yellow | 2024-03-21 20:38:01 | 16 | Region A |
| Red | 2024-03-21 20:38:02 | 204 | Region A |
| Blue | 2024-03-21 20:38:02 | 255 | Region A |
| Green | 2024-03-21 20:38:02 | 213 | Region A |
| Yellow | 2024-03-21 20:38:02 | 18 | Region A |
| Red | 2024-03-21 20:38:03 | 180 | Region A |
| Blue | 2024-03-21 20:38:03 | 254 | Region A |
| Green | 2024-03-21 20:38:03 | 214 | Region A |
| Yellow | 2024-03-21 20:38:03 | 13 | Region A |
| Red | 2024-03-21 20:38:04 | 182 | Region A |
| Blue | 2024-03-21 20:38:04 | 254 | Region A |
| Green | 2024-03-21 20:38:04 | 213 | Region A |
| Yellow | 2024-03-21 20:38:04 | 12 | Region A |
现有代码(逐行处理,不确定是否正确)
我之前没用过Python和scikit-learn,写了这段初步代码,目前还是逐行处理数据:
import pandas as pd from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split from sklearn.preprocessing import LabelEncoder data = pd.read_csv("data/combined.csv", header=0) label_encoder = LabelEncoder() print(data.columns) data["name_encoded"] = label_encoder.fit_transform(data["name"]) data["location_encoded"] = label_encoder.fit_transform(data["location"]) x = data[["rssi", "name_encoded"]] # Features (rssi and encoded name) y = data["location_encoded"] # Target (encoded location) # Split data into training and testing sets x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42) # Create the Random Forest model model = RandomForestClassifier(n_estimators=100) # Train the model model.fit(x_train, y_train) # Make predictions on the testing set y_pred = model.predict(x_test) # Decode predictions location_decoder = LabelEncoder() location_decoder.fit(data["location"]) # Fit the decoder with original locations predicted_locations = location_decoder.inverse_transform(y_pred) print("Predicted locations:", predicted_locations)
内容的提问来源于stack exchange,提问作者senfo
相关产品推荐
相关产品推荐

