You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术问询:如何从给定文本提取地址详情及不含 的目标地址

Hey there! Let's break down your two address extraction questions step by step with practical, actionable solutions.

Address Extraction Solutions for Your Text

1. How to Extract Address Details from Given Text?

When working with text that has clear section labels like Address :, you can combine regex pattern matching and structured parsing to reliably pull out full address blocks. Here's a concrete approach using Python:

Core Strategy:

  • Locate lines starting with Address : to mark the start of each address.
  • Capture all subsequent lines until you hit the next section header (like Name :, Entity, or Full Particulars of Remittance).

Code Example:

import re

# Your input text
raw_text = """From :
Name : NAMITA ROY
Address : 39/4B
 GOPALNAGAR ROAD
 ALIPORE
 KOLKATA,WEST BENGAL
 700027
Entity 
Name : SWARNABARSA PROJECTS PRIVATE LIMITED
Address : 90A
 RAJ SEKHAR BOSE SARANI, FLAT NO.1D, 1ST FLOOR
 KOLKATA,WEST BENGAL
 INDIA - 700025
Full Particulars of Remittance
Service Type: eFiling
"""

# Regex to capture address blocks after "Address :"
address_regex = re.compile(r'Address :\s*(.*?)(?=\n(?:Name|Entity|Full Particulars|Service Type):?)', re.DOTALL)
raw_addresses = address_regex.findall(raw_text)

# Clean up line breaks and extra spaces
clean_addresses = [addr.strip().replace('\n', ' ') for addr in raw_addresses]

# Output results
print("Extracted Addresses:")
for i, addr in enumerate(clean_addresses, 1):
    print(f"{i}. {addr}")

Output:

Extracted Addresses:
1. 39/4B GOPALNAGAR ROAD ALIPORE KOLKATA,WEST BENGAL 700027
2. 90A RAJ SEKHAR BOSE SARANI, FLAT NO.1D, 1ST FLOOR KOLKATA,WEST BENGAL INDIA - 700025

The regex uses a positive lookahead to stop capturing when it detects the next section marker, ensuring you get the full address without extra unrelated content.


2. How to Extract the Address Excluding the "39/4B..." One (Target: 700025)?

To isolate the address ending with 700025 and exclude the 39/4B entry, you can filter the cleaned addresses using unique identifiers like the postal code or associated entity name. Here's how to adjust the code:

Code Example:

# Filter for the address containing 700025 and excluding 39/4B
target_address = next(addr for addr in clean_addresses if "700025" in addr and "39/4B" not in addr)

print("\nTarget Address (700025):")
print(target_address)

Output:

Target Address (700025):
90A RAJ SEKHAR BOSE SARANI, FLAT NO.1D, 1ST FLOOR KOLKATA,WEST BENGAL INDIA - 700025

If you prefer not to rely on the postal code, you could also map addresses to their associated entity names first (e.g., link the address to SWARNABARSA PROJECTS PRIVATE LIMITED) then select the corresponding entry.


Quick Tips for Edge Cases:

  • If your text has inconsistent formatting, tweak the regex pattern to handle varying whitespace or section headers.
  • For more robust extraction (even without explicit labels), use NLP libraries like spaCy with pre-trained named entity recognition (NER) models—these can identify addresses as structured entities automatically.

内容的提问来源于stack exchange,提问作者kroy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:06:45