使用Python解析OSM XML提取含bridge标签的way节点至CSV
Fixing Your OSM XML to CSV Bridge Filtering Code
Hey there! Let's break down why you're hitting that AttributeError and get your code working to extract bridge ways into a CSV like you need.
What's Causing the Error?
Your current code uses way.find('.//tag') which only grabs the first tag inside a way. Here's the double problem:
- If a way has no tags at all, this returns
None, and trying to accesstag.attribimmediately throws anAttributeError. - Even if there are tags, you're only checking the first one—so you'll miss ways where the
bridgetag isn't the first in the list.
Corrected Code to Extract Bridge Ways & Save to CSV
This code will handle your large OSM file safely, filter out only the ways with bridge=yes tags, and save all their complete details to a CSV:
import xml.etree.ElementTree as ET import csv # Define columns for our CSV (adjust if you need extra fields) csv_columns = [ 'way_id', 'visible', 'version', 'changeset', 'timestamp', 'user', 'uid', 'nd_refs', 'all_tags' ] # Set up the CSV file for writing with open('bridge_ways.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.DictWriter(csvfile, fieldnames=csv_columns) writer.writeheader() # Iterate through XML efficiently (critical for large files to save memory) for event, elem in ET.iterparse('MER.xml', events=('start', 'end')): if event == 'end' and elem.tag == 'way': # Check if this way has a bridge=yes tag is_bridge = False tag_dict = {} for tag in elem.findall('tag'): tag_key = tag.attrib.get('k') tag_val = tag.attrib.get('v') tag_dict[tag_key] = tag_val if tag_key == 'bridge' and tag_val == 'yes': is_bridge = True if is_bridge: # Collect all data from the matching way way_details = { 'way_id': elem.attrib.get('id'), 'visible': elem.attrib.get('visible'), 'version': elem.attrib.get('version'), 'changeset': elem.attrib.get('changeset'), 'timestamp': elem.attrib.get('timestamp'), 'user': elem.attrib.get('user'), 'uid': elem.attrib.get('uid'), # Combine node refs into a comma-separated string 'nd_refs': ','.join([nd.attrib.get('ref') for nd in elem.findall('nd')]), # Format all tags into a readable string 'all_tags': ','.join([f"{k}:{v}" for k, v in tag_dict.items()]) } writer.writerow(way_details) # Clear element from memory to avoid bloating with large files elem.clear() print("Done! All bridge ways have been saved to bridge_ways.csv")
Key Improvements:
- Safe Tag Checking: We loop through all tags in a way, so we never miss
bridgetags that aren't first, and avoidNone-related errors. - Memory Efficiency: Using
ET.iterparse()and clearing elements after processing keeps memory usage low—perfect for your 4000-way OSM file. - Complete Way Data: The CSV includes every core attribute of the way, all node references, and all associated tags, so you get the full way information as requested.
- Error Resilient: Using
.get()for attribute access prevents crashes if an optional attribute is missing from a way.
Example CSV Output Snippet:
way_id,visible,version,changeset,timestamp,user,uid,nd_refs,all_tags
108534076,true,1,7866393,2011-04-15T02:42:51Z,richlv,47892,1245024935,1245025038,bridge:yes,highway:service
内容的提问来源于stack exchange,提问作者John Smith
相关产品推荐
相关产品推荐

