You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于匹配数据生成对应列,Pandas实现站点线路二元统计

原始数据预处理

你提供的原始代码重复嵌套了pd.DataFrame,先修正并添加列名方便后续处理:

import pandas as pd
df = pd.DataFrame([
    ('Bus Route A', 'Beach Station'),
    ('Bus Route B', 'Village Hotel '),
    ('Bus Route A', 'Amara Sanctuary Resort'),
    ('Bus Route C', 'Village Hotel '),
    ('Bus Route B', 'Beach Station'),
    ('Bus Route C', 'Beach Station')
], columns=['公交线路', '站点名称'])

站点和线路对应关系参考:
站点线路对应关系示意图


功能实现代码

# 交叉统计每个站点在各线路的出现情况,自动去重标记1/0
station_route_matrix = pd.crosstab(df['站点名称'], df['公交线路']).clip(upper=1).reset_index()

# 转换为要求的二元格式元组列表
binary_stat = [tuple(row) for row in station_route_matrix.values]

# 输出结果验证
print(binary_stat)

输出结果

运行后得到的binary_stat完全符合你要求的结构:

[('Amara Sanctuary Resort', 1, 0, 0), ('Beach Station', 1, 1, 1), ('Village Hotel ', 0, 1, 1)]

生成的station_route_matrix已经自带三个公交线路的统计列,可直接用于后续计算。


内容的提问来源于stack exchange,提问作者Sahil Kamboj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 01:24:04