You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas处理OpenStreetMap .pbf文件时遇空DataFrame及NameError问题

Why You're Getting an Empty DataFrame When Reading OSM .pbf Files

Let's break down exactly what's going wrong here and how to fix it step by step:

1. You're Accidentally Overwriting Your Loaded Data

Looking at your code snippet:

df_osm = pd.DataFrame(handler.osm_data, columns=data_colnames)
tag_genome = pd.DataFrame(columns=data_colnames)
df_osm = tag_genome.sort_values(by=['type', 'id', 'ts'])

You first create df_osm from the data collected by your handler, but then immediately overwrite it with a sorted version of tag_genome—which you initialized as an empty DataFrame! That's why you end up with an empty result, even if handler.osm_data had valid data in it.

2. Your Handler Might Not Be Properly Parsing the .pbf File

If handler.osm_data was already empty to begin with, fixing the overwriting issue won't help. Most OSM .pbf parsing tools (like pyosmium or osmnx) require you to define a handler class that explicitly collects data from nodes, ways, and relations as it parses the file. If your handler doesn't have the right callback methods set up, it won't populate osm_data at all.

Fixed Example Code

Here's a corrected workflow using pyosmium (a popular library for parsing .pbf files) that properly collects data and avoids overwriting your loaded content:

import osmium
import pandas as pd

# Define a handler to collect OSM tag data
class OSMDataHandler(osmium.SimpleHandler):
    def __init__(self):
        super().__init__()
        self.osm_data = []  # This will store our parsed data

    # Helper method to record tag details for any OSM element
    def _log_tag(self, elem, elem_type):
        for tag in elem.tags:
            self.osm_data.append({
                'type': elem_type,
                'id': elem.id,
                'version': elem.version,
                'visible': elem.visible,
                'ts': elem.timestamp,
                'uid': elem.uid,
                'user': elem.user,
                'chgset': elem.changeset,
                'ntags': len(elem.tags),
                'tagkey': tag.k,
                'tagvalue': tag.v
            })

    # Override methods to handle nodes, ways, and relations
    def node(self, n):
        self._log_tag(n, 'node')

    def way(self, w):
        self._log_tag(w, 'way')

    def relation(self, r):
        self._log_tag(r, 'relation')

# Parse your .pbf file
handler = OSMDataHandler()
handler.apply_file('your_file.pbf', locations=False)  # Set locations=True if you need geocoordinates

# Create DataFrames without overwriting data
data_colnames = ['type', 'id', 'version', 'visible', 'ts', 'uid', 'user', 'chgset', 'ntags', 'tagkey', 'tagvalue']
df_osm = pd.DataFrame(handler.osm_data, columns=data_colnames)

# If you need tag_genome, use a copy of df_osm instead of an empty frame
tag_genome = df_osm.copy()
df_osm_sorted = tag_genome.sort_values(by=['type', 'id', 'ts'])

print(df_osm_sorted.head())

Quick Verification Checks

  • Before creating the DataFrame, run print(len(handler.osm_data))—if it returns 0, your handler isn't collecting data. Double-check that you've implemented the node, way, and relation methods correctly.
  • Confirm your .pbf file isn't corrupted or empty: use a tool like osmium-tool to inspect it with this command:
    osmium fileinfo your_file.pbf
    

内容的提问来源于stack exchange,提问作者user8531240

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:51:11