You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Pandas实现FullAddress多列匹配生成带标签地址列

Pandas 地址字段标签化解决方案

原始数据结构

PremiseThoroughfareLocalityPostalCodeCountryFullAddress
Yew Tree LaneHolmbridgeHD9 2NRN IrelandOld Thorn, Yew Tree Lane, Holmbridge HD9 2NR, N Ireland
3Cysgod Y CastellLlandudno JunctionLL31 9LJUk3 Cysgod Y Castell, Llandudno Junction LL31 9LJ
1168Christchurch RoadBournemouthBH7 6DYWales UK1168 Christchurch Road, BH7 6DY Bournemouth

目标输出结构

FullAddressFullAdressWithTag
Old Thorn, Yew Tree Lane, Holmbridge HD9 2NR, N IrelandOld^Others Thorn^Others, Yew^Thoroughfare Tree^Thoroughfare Lane^Thoroughfare, Holmbridge^Locality HD9^PostalCode 2NR^PostalCode, N^Country Ireland^Country
3 Cysgod Y Castell, Llandudno Junction LL31 9LJ3^Premise Cysgod^Thoroughfare Y^Thoroughfare Castell^Thoroughfare, Llandudno^Locality Junction^Locality LL31^PostalCode 9LJ^PostalCode
1168 Christchurch Road, BH7 6DY Bournemouth1168^Premise Christchurch^Thoroughfare Road^Thoroughfare, BH7^PostalCode 6DY^PostalCode Bournemouth^Locality

规则说明

  • 将FullAddress中的每个单词与Premise、Thoroughfare、Locality、PostalCode、Country列的内容匹配
  • 匹配到对应列的单词,格式化为单词^列名;未匹配到的标记为单词^Others
  • 支持FullAddress字段的任意排列顺序,需适配百万级数据集的高效处理

实现代码

import pandas as pd

def tag_address(row):
    # 按短语长度降序处理,优先匹配长字段内容
    fields = sorted(
        [("Premise", row["Premise"]), ("Thoroughfare", row["Thoroughfare"]), 
         ("Locality", row["Locality"]), ("PostalCode", row["PostalCode"]), 
         ("Country", row["Country"])],
        key=lambda x: len(str(x[1]).split()), reverse=True
    )
    
    tag_map = {}
    for tag, value in fields:
        if pd.notna(value) and value.strip():
            for word in str(value).split():
                if word not in tag_map:
                    tag_map[word] = tag
    
    # 拆分地址并添加标签
    address_segments = row["FullAddress"].split(", ")
    tagged_segments = []
    for seg in address_segments:
        tagged_words = [f"{word}^{tag_map.get(word, 'Others')}" for word in seg.split()]
        tagged_segments.append(" ".join(tagged_words))
    
    return ", ".join(tagged_segments)

# 应用函数生成标签列
df["FullAdressWithTag"] = df.apply(tag_address, axis=1)
# 提取结果数据集
result_df = df[["FullAddress", "FullAdressWithTag"]].copy()

代码说明

  1. 长短语优先匹配:按字段内容的单词数量降序处理,确保长短语中的单词优先被标记对应列标签,避免短字段内容的单词覆盖
  2. 去重标签分配:每个单词仅保留第一个匹配到的标签,保证一致性
  3. 结构保留:拆分FullAddress的逗号分隔段后再处理单词,最后重组,完全保留原始地址的标点与分段结构
  4. 空值处理:自动跳过空值或空字符串的字段,避免无效匹配

内容的提问来源于stack exchange,提问作者Puteri Hasya Damia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:35:21