基于Pandas实现FullAddress多列匹配生成带标签地址列
Pandas 地址字段标签化解决方案
原始数据结构
| Premise | Thoroughfare | Locality | PostalCode | Country | FullAddress |
|---|---|---|---|---|---|
| Yew Tree Lane | Holmbridge | HD9 2NR | N Ireland | Old Thorn, Yew Tree Lane, Holmbridge HD9 2NR, N Ireland | |
| 3 | Cysgod Y Castell | Llandudno Junction | LL31 9LJ | Uk | 3 Cysgod Y Castell, Llandudno Junction LL31 9LJ |
| 1168 | Christchurch Road | Bournemouth | BH7 6DY | Wales UK | 1168 Christchurch Road, BH7 6DY Bournemouth |
目标输出结构
| FullAddress | FullAdressWithTag |
|---|---|
| Old Thorn, Yew Tree Lane, Holmbridge HD9 2NR, N Ireland | Old^Others Thorn^Others, Yew^Thoroughfare Tree^Thoroughfare Lane^Thoroughfare, Holmbridge^Locality HD9^PostalCode 2NR^PostalCode, N^Country Ireland^Country |
| 3 Cysgod Y Castell, Llandudno Junction LL31 9LJ | 3^Premise Cysgod^Thoroughfare Y^Thoroughfare Castell^Thoroughfare, Llandudno^Locality Junction^Locality LL31^PostalCode 9LJ^PostalCode |
| 1168 Christchurch Road, BH7 6DY Bournemouth | 1168^Premise Christchurch^Thoroughfare Road^Thoroughfare, BH7^PostalCode 6DY^PostalCode Bournemouth^Locality |
规则说明
- 将
FullAddress中的每个单词与Premise、Thoroughfare、Locality、PostalCode、Country列的内容匹配 - 匹配到对应列的单词,格式化为
单词^列名;未匹配到的标记为单词^Others - 支持
FullAddress字段的任意排列顺序,需适配百万级数据集的高效处理
实现代码
import pandas as pd def tag_address(row): # 按短语长度降序处理,优先匹配长字段内容 fields = sorted( [("Premise", row["Premise"]), ("Thoroughfare", row["Thoroughfare"]), ("Locality", row["Locality"]), ("PostalCode", row["PostalCode"]), ("Country", row["Country"])], key=lambda x: len(str(x[1]).split()), reverse=True ) tag_map = {} for tag, value in fields: if pd.notna(value) and value.strip(): for word in str(value).split(): if word not in tag_map: tag_map[word] = tag # 拆分地址并添加标签 address_segments = row["FullAddress"].split(", ") tagged_segments = [] for seg in address_segments: tagged_words = [f"{word}^{tag_map.get(word, 'Others')}" for word in seg.split()] tagged_segments.append(" ".join(tagged_words)) return ", ".join(tagged_segments) # 应用函数生成标签列 df["FullAdressWithTag"] = df.apply(tag_address, axis=1) # 提取结果数据集 result_df = df[["FullAddress", "FullAdressWithTag"]].copy()
代码说明
- 长短语优先匹配:按字段内容的单词数量降序处理,确保长短语中的单词优先被标记对应列标签,避免短字段内容的单词覆盖
- 去重标签分配:每个单词仅保留第一个匹配到的标签,保证一致性
- 结构保留:拆分
FullAddress的逗号分隔段后再处理单词,最后重组,完全保留原始地址的标点与分段结构 - 空值处理:自动跳过空值或空字符串的字段,避免无效匹配
内容的提问来源于stack exchange,提问作者Puteri Hasya Damia
相关产品推荐
相关产品推荐

