You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中避免OpenStreetMap数据的商铺重复统计

解决OSM商铺重复统计的方案

你遇到的问题本质是OSM中大型商铺的多节点映射,以及单纯靠名称/坐标取整去重的局限性,以下是几个可行的解决思路:

1. 扩展处理OSM的Way和Relation元素

大型商铺(如百货商场)通常不是用单个Node表示,而是用**Way(多边形区域)或Relation(建筑群集合)**来标记,这些元素会包含shop标签,同时关联多个Node。你的代码目前只处理Node,导致遗漏了这些核心数据,还会把同一商铺的附属Node当成独立商铺。

修改Handler,同时处理Node、Way、Relation:

import osmium
import pandas as pd

class ShopHandler(osmium.SimpleHandler):
    def __init__(self):
        super().__init__()
        self.shops = {}

    def _process_shop(self, elem, elem_type):
        if 'shop' not in elem.tags:
            return
        # 优先用唯一标识作为键,比如wikidata、ref
        unique_key = elem.tags.get('wikidata') or elem.tags.get('ref') or elem.tags.get('name', 'Unnamed')
        # 补充地址信息增强唯一性
        addr = f"{elem.tags.get('addr:street','')} {elem.tags.get('addr:housenumber','')}".strip()
        if addr:
            unique_key += f"_{addr}"
        
        if unique_key not in self.shops:
            # 对于Way/Relation,取包围盒中心坐标
            if elem_type == 'node':
                lat, lon = elem.location.lat, elem.location.lon
            else:
                bounds = elem.envelope()
                lat = (bounds.top + bounds.bottom)/2
                lon = (bounds.left + bounds.right)/2
            self.shops[unique_key] = {
                'elem_ids': [elem.id],
                'elem_type': elem_type,
                'lat': lat,
                'lon': lon,
                'shop_type': elem.tags.get('shop', 'Unknown'),
                'name': elem.tags.get('name', 'Unnamed'),
                'addr': addr
            }
        else:
            self.shops[unique_key]['elem_ids'].append(elem.id)

    def node(self, n):
        self._process_shop(n, 'node')

    def way(self, w):
        self._process_shop(w, 'way')

    def relation(self, r):
        self._process_shop(r, 'relation')

handler = ShopHandler()
handler.apply_file(file_path)
shops_df = pd.DataFrame.from_dict(handler.shops, orient='index').reset_index(drop=True)

2. 用空间聚类处理无唯一标识的节点

如果某些商铺没有wikidata/ref这类标签,用DBSCAN密度聚类来识别邻近的同商铺节点——它会根据距离和密度自动分组,不会像坐标取整那样误合并邻近商铺。

示例代码(基于已收集的节点数据):

from sklearn.cluster import DBSCAN
import numpy as np

# 假设shops_df里有lat、lon、name、shop_type列
coords = shops_df[['lat', 'lon']].values
# 转换坐标到米(用Haversine公式,DBSCAN的eps单位设为米)
kms_per_radian = 6371.0088
epsilon = 0.05  # 50米阈值,可根据商铺大小调整
db = DBSCAN(eps=epsilon/kms_per_radian, min_samples=1, algorithm='ball_tree', metric='haversine').fit(np.radians(coords))

# 添加聚类标签,同一聚类视为同一商铺
shops_df['cluster_id'] = db.labels_
# 按聚类和商铺类型统计独立数量
unique_shops = shops_df.groupby(['cluster_id', 'shop_type'])['name'].first().reset_index()
shop_type_count = unique_shops['shop_type'].value_counts()

3. 结合多维度标签去重

除了唯一标识和空间聚类,还可以结合brand(品牌)、operator(运营商)、addr:postcode(邮编)等标签,构建更精准的去重键:

# 构建复合去重键
unique_key = f"{elem.tags.get('brand','')}_{elem.tags.get('name','')}_{elem.tags.get('addr:postcode','')}_{elem.tags.get('addr:housenumber','')}".strip()

内容的提问来源于stack exchange,提问作者gutmiz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 00:47:52