使用Pandas处理OpenStreetMap .pbf文件时遇空DataFrame及NameError问题
Let's break down exactly what's going wrong here and how to fix it step by step:
1. You're Accidentally Overwriting Your Loaded Data
Looking at your code snippet:
df_osm = pd.DataFrame(handler.osm_data, columns=data_colnames) tag_genome = pd.DataFrame(columns=data_colnames) df_osm = tag_genome.sort_values(by=['type', 'id', 'ts'])
You first create df_osm from the data collected by your handler, but then immediately overwrite it with a sorted version of tag_genome—which you initialized as an empty DataFrame! That's why you end up with an empty result, even if handler.osm_data had valid data in it.
2. Your Handler Might Not Be Properly Parsing the .pbf File
If handler.osm_data was already empty to begin with, fixing the overwriting issue won't help. Most OSM .pbf parsing tools (like pyosmium or osmnx) require you to define a handler class that explicitly collects data from nodes, ways, and relations as it parses the file. If your handler doesn't have the right callback methods set up, it won't populate osm_data at all.
Fixed Example Code
Here's a corrected workflow using pyosmium (a popular library for parsing .pbf files) that properly collects data and avoids overwriting your loaded content:
import osmium import pandas as pd # Define a handler to collect OSM tag data class OSMDataHandler(osmium.SimpleHandler): def __init__(self): super().__init__() self.osm_data = [] # This will store our parsed data # Helper method to record tag details for any OSM element def _log_tag(self, elem, elem_type): for tag in elem.tags: self.osm_data.append({ 'type': elem_type, 'id': elem.id, 'version': elem.version, 'visible': elem.visible, 'ts': elem.timestamp, 'uid': elem.uid, 'user': elem.user, 'chgset': elem.changeset, 'ntags': len(elem.tags), 'tagkey': tag.k, 'tagvalue': tag.v }) # Override methods to handle nodes, ways, and relations def node(self, n): self._log_tag(n, 'node') def way(self, w): self._log_tag(w, 'way') def relation(self, r): self._log_tag(r, 'relation') # Parse your .pbf file handler = OSMDataHandler() handler.apply_file('your_file.pbf', locations=False) # Set locations=True if you need geocoordinates # Create DataFrames without overwriting data data_colnames = ['type', 'id', 'version', 'visible', 'ts', 'uid', 'user', 'chgset', 'ntags', 'tagkey', 'tagvalue'] df_osm = pd.DataFrame(handler.osm_data, columns=data_colnames) # If you need tag_genome, use a copy of df_osm instead of an empty frame tag_genome = df_osm.copy() df_osm_sorted = tag_genome.sort_values(by=['type', 'id', 'ts']) print(df_osm_sorted.head())
Quick Verification Checks
- Before creating the DataFrame, run
print(len(handler.osm_data))—if it returns 0, your handler isn't collecting data. Double-check that you've implemented thenode,way, andrelationmethods correctly. - Confirm your .pbf file isn't corrupted or empty: use a tool like
osmium-toolto inspect it with this command:osmium fileinfo your_file.pbf
内容的提问来源于stack exchange,提问作者user8531240

