You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为跨国城市列表及非标准化地址补充地理属性字段?

一、为多国城市列表补充地理信息

可行方案:

  • 使用离线结构化地理数据库:GeoNames是免费覆盖全球的地理数据集,包含城市、市政、州/省、邮编等全量字段。下载对应国家的数据集后,通过城市名(注意匹配多语言别名)关联匹配,可批量补充缺失信息,无需调用外部接口,适合大规模数据处理。
  • 批量调用免费地理编码接口:用Python的geopy库对接OpenStreetMap的Nominatim接口,通过「城市+国家」的组合查询获取结构化地址详情。示例代码:
from geopy.geocoders import Nominatim
geolocator = Nominatim(user_agent="your_project_label")

def fetch_geo_details(city, country):
    location = geolocator.geocode(f"{city}, {country}", addressdetails=True)
    if not location:
        return ("NA", "NA", "NA")
    addr = location.raw.get("address", {})
    return (
        addr.get("municipality", "NA"),
        addr.get("state", "NA"),
        addr.get("region", "NA")
    )

# 调用示例
municipality, state, region = fetch_geo_details("Zürich", "Switzerland")
  • 多语言兼容处理:阿拉伯语、葡萄牙语等非拉丁字符的城市名,GeoNames和Nominatim均可正常识别,匹配时保留原始字符避免转码错误。
二、拆分格式不统一的居住地数据字段

核心思路:先拆分国家,再按规则匹配各级地理单元,最后补全城乡属性

  1. 分离国家字段:所有记录的国家均位于最后一个逗号之后,先按最后一个逗号分割,分离出国家,剩余部分处理城市/州等信息。
  2. 正则规则批量匹配地理单元:针对不同格式写针对性规则,用Python的re和pandas批量处理,示例代码:
import re
import pandas as pd

def parse_residence(address):
    # 分割国家与地址主体
    split_parts = [p.strip() for p in address.split(',')]
    country = split_parts[-1]
    address_body = ','.join(split_parts[:-1])
    
    # 初始化字段为NA
    city = "NA"
    municipality = "NA"
    district = "NA"
    state = "NA"
    rural_urban = "NA"
    
    # 匹配带"State of"的格式(如巴西案例)
    state_match = re.search(r'- State of ([\w\s]+)', address_body)
    if state_match:
        state = state_match.group(1).strip()
        city = address_body.split('-')[0].strip()
    # 匹配城市+州的格式(如印度案例)
    elif len(address_body.split(',')) >= 2:
        city = address_body.split(',')[0].strip()
        state = address_body.split(',')[1].strip()
    # 单城市格式(如瑞士、巴林案例)
    else:
        city = address_body.strip()
    
    # 识别城乡属性:通过关键词匹配
    if re.search(r'village|rural|countryside', address_body, re.IGNORECASE):
        rural_urban = "Rural"
    elif re.search(r'city|town|مدينة', address_body, re.IGNORECASE):
        rural_urban = "Urban"
    
    return pd.Series([city, municipality, district, state, country, rural_urban])

# 测试示例数据
sample_data = pd.DataFrame({
    "residence": [
        "مدينة حمد، Bahrain", 
        "Indore, Madhya Pradesh, India", 
        "Zürich, Switzerland", 
        "São Luís - State of Maranhão, Brazil"
    ]
})

# 生成拆分后的字段
sample_data[["城市/乡村名称", "市政", "区县", "州", "国家", "城乡属性"]] = sample_data["residence"].apply(parse_residence)
  1. 处理特殊格式(邮编+城市):对于带邮编的记录,先用正则提取邮编(re.findall(r'\d+', address_body)),剩余部分作为城市名,再通过地理编码接口反查对应的州/市政信息补全字段。
  2. 手动校验补全:筛选出规则未覆盖的模糊记录,手动核对补全,确保数据准确性。

内容的提问来源于stack exchange,提问作者Olivia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 10:43:11