如何用dc.js绘制非线性数据?嵌套JSON的Crossfilter分组方案
处理嵌套/非线性JSON患者数据:Crossfilter + dc.js 实现多维度筛选
看起来你遇到了Crossfilter处理嵌套数组型数据的常见问题——Crossfilter本质上是为扁平表格数据设计的,嵌套的数组(比如problems、measurements)和缺失字段会让分组筛选变得棘手。不过别担心,核心思路是先把非线性数据扁平化,再基于扁平数据构建Crossfilter维度和dc.js图表,就能实现你要的多维度联动筛选。
1. 先把嵌套数据扁平化
每个患者的problems、measurements是数组,意味着一个患者可能对应多个疾病或测量值。我们需要把这些数组展开,让每个条目(比如每个疾病)对应一条包含患者核心信息的记录——这样Crossfilter才能正确识别并分组这些维度。
举个具体的扁平化函数,针对你的数据结构:
function flattenPatientData(patients) { const flattenedRecords = []; patients.forEach(patient => { // 提取患者基础信息 const baseInfo = { patient_id: patient.patient_id, gender: patient.demographics?.gender || 'unknown', age: patient.demographics?.age || null }; // 处理measurements:这里提取心率数据(如果有) const pulseData = patient.measurements?.find(m => m.kind === 'pulse'); const measurementInfo = { pulse_value: pulseData?.value ? parseInt(pulseData.value) : null, pulse_date: pulseData?.measurementDate || null }; // 处理problems数组:每个疾病生成一条记录 if (patient.problems && patient.problems.length > 0) { patient.problems.forEach(problem => { flattenedRecords.push({ ...baseInfo, ...measurementInfo, problem_name: problem.name_title, problem_category: problem.category, problem_startDate: problem.startDate }); }); } else { // 没有疾病的患者,单独生成一条记录(避免丢失数据) flattenedRecords.push({ ...baseInfo, ...measurementInfo, problem_name: '无记录', problem_category: null, problem_startDate: null }); } }); return flattenedRecords; }
调用这个函数后,你的嵌套数据就会变成每条记录对应「一位患者+一个疾病」的扁平结构,完美适配Crossfilter的需求。
2. 构建Crossfilter维度和分组
基于扁平数据,我们可以为每个需要筛选的字段创建维度,同时根据图表需求定义分组逻辑(比如按年龄区间、疾病名称分组)。
// 1. 先扁平化原始数据 const rawPatientData = [/* 你的原始JSON数据 */]; const flattenedData = flattenPatientData(rawPatientData); // 2. 创建Crossfilter实例 const cf = crossfilter(flattenedData); // --- 定义维度和分组 --- // 疾病维度:用于筛选特定疾病 const problemDimension = cf.dimension(d => d.problem_name); // 统计每个疾病对应的**唯一患者数**(避免重复计数) const problemUniqueGroup = problemDimension.group().reduce( (acc, record) => { acc[record.patient_id] = true; return acc; }, (acc, record) => { delete acc[record.patient_id]; return acc; }, () => ({}) ).reduce((total, patientMap) => Object.keys(patientMap).length, 0); // 年龄维度:用于年龄分布图表 const ageDimension = cf.dimension(d => d.age); // 按10岁区间分组 const ageGroup = ageDimension.group(age => Math.floor(age / 10) * 10); // 性别维度:用于性别分布 const genderDimension = cf.dimension(d => d.gender); const genderGroup = genderDimension.group().reduceCount(); // 心率维度:用于心率分布 const pulseDimension = cf.dimension(d => d.pulse_value); // 按10bpm区间分组 const pulseGroup = pulseDimension.group(pulse => pulse ? Math.floor(pulse / 10) * 10 : null);
这里重点注意:如果直接用reduceCount,会因为一个患者对应多条疾病记录而重复计数,所以我们用自定义的reduce逻辑来统计唯一患者ID的数量,这样结果更准确。
3. 用dc.js创建联动图表
现在可以基于这些维度和分组创建dc.js图表,它们会自动联动筛选——比如点击疾病图表中的某个疾病,年龄、性别、心率图表会自动更新为该疾病患者的对应数据。
// 疾病柱状图 const problemChart = dc.barChart('#problem-chart-container'); problemChart .dimension(problemDimension) .group(problemUniqueGroup) .x(d3.scaleBand()) .xUnits(dc.units.ordinal) .xAxisLabel('疾病名称') .yAxisLabel('患者数量') .width(600) .height(300); // 年龄直方图 const ageChart = dc.barChart('#age-chart-container'); ageChart .dimension(ageDimension) .group(ageGroup) .x(d3.scaleLinear().domain([0, 100])) .xAxisLabel('年龄区间(岁)') .yAxisLabel('患者数量') .width(600) .height(300); // 性别饼图 const genderChart = dc.pieChart('#gender-chart-container'); genderChart .dimension(genderDimension) .group(genderGroup) .radius(100) .width(300) .height(300); // 渲染所有图表 dc.renderAll();
关键注意事项
- 缺失值处理:给缺失的字段(比如无疾病、无测量值的患者)设置默认值(如
null、'无记录'),避免Crossfilter或dc.js报错。 - 多数组展开:如果需要同时基于
measurements和problems筛选,可能需要进一步扁平化(比如每个测量值+每个疾病对应一条记录),但要根据你的业务需求权衡数据冗余度。 - 性能优化:如果患者数据量很大,扁平化后记录数会翻倍,建议在前端处理前先做数据分页或后端预处理。
内容的提问来源于stack exchange,提问作者galatia
相关产品推荐
相关产品推荐

