You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用WEKA处理洛杉矶犯罪数据时,构建可解析CSV文件遇难题

洛杉矶犯罪数据适配WEKA的CSV构建方案

一、导出指定字段的原始CSV

  • 打开数据源页面,找到导出按钮(一般标着Export)
  • 选择CSV格式,在自定义字段列表里勾选AREA name、Vict age、Vict descent、Vict sex、Date Occ,确认后导出文件。

二、预处理CSV适配WEKA要求

WEKA对输入格式有严格要求,按以下步骤调整:

1. 修正表头命名

WEKA不接受表头带空格,直接把表头改成:

  • AREA name → AREA_name
  • Vict age → Vict_age
  • Vict descent → Vict_descent
  • Vict sex → Vict_sex
  • Date Occ → Date_Occ

2. 处理缺失值

WEKA用?标记缺失值,批量替换CSV里的空值:

  • Vict_age:空值直接换成?;如果有非数字内容,要么删行,要么统一替换成?
  • Vict_sex/Vict_descent:空值替换为?,同时统一分类值的大小写(比如把f改成F,避免WEKA识别成不同类别)

3. 格式化日期字段

WEKA需要明确的日期格式,把Date_Occ转换成yyyy-MM-dd HH:mm:ss格式:

  • 用Excel的话,选中日期列,设置单元格格式为yyyy-mm-dd hh:mm:ss
  • 用Python批量处理的话,代码如下:
    import pandas as pd
    df = pd.read_csv("你的导出文件.csv")
    df['Date_Occ'] = pd.to_datetime(df['Date_Occ']).dt.strftime('%Y-%m-%d %H:%M:%S')
    df.to_csv("weka_crime_data.csv", index=False)
    

4. 可选:添加WEKA属性声明

如果想让WEKA自动识别字段类型,可以在CSV第一行前插入以下内容:

@relation CrimeData
@attribute AREA_name string
@attribute Vict_age numeric
@attribute Vict_descent string
@attribute Vict_sex string
@attribute Date_Occ date "yyyy-MM-dd HH:mm:ss"
@data

然后把原始数据行接在后面即可。

三、导入WEKA验证

  • 打开WEKA Explorer,点击Open file选中处理后的CSV
  • 切换到Preprocess标签,检查字段类型是否正确:Vict_age是数值型,其他为字符串/日期型
  • 要是有异常值(比如年龄超过120),用WEKA的Filter工具(比如RemoveWithValues)清理就行

内容的提问来源于stack exchange,提问作者Frederik N. Pedersen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 22:42:50