相同代码在.ipynb正常运行但在.py文件中报错的问题求助
解决OSMnx加载.osm文件时的UnicodeDecodeError问题
问题本质
Windows命令行环境下,Python默认采用cp1252编码打开文件,但Geofabrik下载的.osm文件是UTF-8编码,导致ElementTree解析时出现解码失败。而Jupyter Notebook会自动以UTF-8编码处理文件,因此无问题。
快速解决方案:设置Python默认编码为UTF-8
在运行脚本前,在命令行设置环境变量:
set PYTHONUTF8=1 python your_script.py
或者在Python脚本开头添加以下代码,全局设置默认编码:
import os os.environ['PYTHONUTF8'] = '1' # 后续正常调用OSMnx import osmnx as ox network = 'network_file.osm' G = ox.graph_from_xml(network)
替代方案:手动处理文件编码(无需全局设置)
如果不想修改全局编码,可以重写OSMnx内部的文件解析逻辑,强制用UTF-8读取文件:
import xml.etree.ElementTree as ET import osmnx as ox from osmnx import settings, osm_xml def _custom_parse_osm_xml(filepath, node_attrs=None, way_attrs=None, rel_attrs=None, node_tags=None, way_tags=None, rel_tags=None): node_attrs = node_attrs or settings.osm_xml_node_attrs way_attrs = way_attrs or settings.osm_xml_way_attrs rel_attrs = rel_attrs or settings.osm_xml_rel_attrs node_tags = node_tags or settings.osm_xml_node_tags way_tags = way_tags or settings.osm_xml_way_tags rel_tags = rel_tags or settings.osm_xml_rel_tags nodes = {} ways = {} relations = {} # 强制用UTF-8编码打开文件 with open(filepath, 'r', encoding='utf-8') as f: for event, elem in ET.iterparse(f): if elem.tag == 'node': node = {'id': elem.attrib['id']} for attr in node_attrs: if attr in elem.attrib: node[attr] = elem.attrib[attr] node['tags'] = {tag.attrib['k']: tag.attrib['v'] for tag in elem.findall('tag') if tag.attrib['k'] in node_tags} nodes[node['id']] = node elif elem.tag == 'way': way = {'id': elem.attrib['id']} for attr in way_attrs: if attr in elem.attrib: way[attr] = elem.attrib[attr] way['nodes'] = [nd.attrib['ref'] for nd in elem.findall('nd')] way['tags'] = {tag.attrib['k']: tag.attrib['v'] for tag in elem.findall('tag') if tag.attrib['k'] in way_tags} ways[way['id']] = way elif elem.tag == 'relation': relation = {'id': elem.attrib['id']} for attr in rel_attrs: if attr in elem.attrib: relation[attr] = elem.attrib[attr] relation['members'] = [{k: member.attrib[k] for k in ['type', 'ref', 'role']} for member in elem.findall('member')] relation['tags'] = {tag.attrib['k']: tag.attrib['v'] for tag in elem.findall('tag') if tag.attrib['k'] in rel_tags} relations[relation['id']] = relation elem.clear() return nodes, ways, relations def custom_overpass_json_from_file(filepath): # 强制用UTF-8编码解析XML根节点 with open(filepath, 'r', encoding='utf-8') as f: root = ET.parse(f).getroot() root_attrs = root.attrib nodes, ways, relations = _custom_parse_osm_xml(filepath) response_json = osm_xml._create_overpass_json(nodes, ways, relations, root_attrs) return response_json # 替换OSMnx原有的解析函数 osm_xml._overpass_json_from_file = custom_overpass_json_from_file # 正常加载.osm文件 network = 'network_file.osm' G = ox.graph_from_xml(network)
为什么之前的ChatGPT方案失败?
ox.graph_from_xml只接受文件路径字符串或字节流路径,不接受已经打开的TextIOWrapper对象,因此直接传入文件句柄会触发TypeError。
内容的提问来源于stack exchange,提问作者David Micallef
相关产品推荐
相关产品推荐

