You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

技术需求:从多结果页面提取指定地址的目标企业数据

定向提取特定地址企业信息的解决方案

没问题,我来帮你搞定这个需求!针对你提供的HTML片段,我们只需要先定位到包含目标地址「311 South Swall Drive」的企业区块,再从中提取指定字段即可,具体实现思路和代码示例如下:

核心逻辑

每个企业的完整信息都包裹在<section class="org">标签中,我们可以:

  1. 遍历所有企业区块
  2. 检查区块内的streetAddress文本是否匹配目标地址
  3. 找到匹配区块后,提取其中的Status、Registration等字段

实现方案(Python + BeautifulSoup)

这是最常用的网页解析方式,代码清晰易懂:

from bs4 import BeautifulSoup

# 假设你的HTML源码存在这个变量里
html_content = """
<!-- 这里放你提供的HTML片段 -->
"""

soup = BeautifulSoup(html_content, 'html.parser')

# 遍历所有企业区块
for org_section in soup.find_all('section', class_='org'):
    # 获取当前企业的街道地址
    street_addr = org_section.find('span', itemprop='streetAddress')
    if street_addr and street_addr.text.strip() == '311 South Swall Drive':
        # 提取需要的字段
        target_info = {}
        # 遍历所有属性行
        for prop in org_section.find_all('p', class_='b-business-item_props'):
            title = prop.find('span', class_='b-business-item_title').text.strip().rstrip(':')
            value = prop.find('span', class_='b-business-item_value').text.strip()
            target_info[title] = value
        
        # 输出结果
        print("目标企业信息:")
        for key, val in target_info.items():
            print(f"- {key}: {val}")
        break  # 找到目标后停止遍历

运行这段代码后,你会得到精准的结果:

  • Status: Inactive
  • Registration: Sep 26, 2006
  • State ID: C2904860
  • Business type: Articles of Incorporation
  • Member: Ashwant Venkatram (President, inactive)

另一种方案:用XPath直接定位(适用于Scrapy、lxml)

如果你用Scrapy或者lxml这类支持XPath的工具,可以直接用一条XPath表达式定位到目标区块,再提取字段:

from lxml import etree

tree = etree.HTML(html_content)

# 定位目标企业区块
target_section = tree.xpath("//section[@class='org'][.//span[@itemprop='streetAddress' and text()='311 South Swall Drive']]")[0]

# 提取字段
status = target_section.xpath(".//span[text()='Status:']/following-sibling::span/text()")[0].strip()
registration = target_section.xpath(".//span[text()='Registration:']/following-sibling::span/text()")[0].strip()
state_id = target_section.xpath(".//span[text()='State ID:']/following-sibling::span/text()")[0].strip()
business_type = target_section.xpath(".//span[text()='Business type:']/following-sibling::span/text()")[0].strip()
member = target_section.xpath(".//span[text()='Member:']/following-sibling::span/text()")[0].strip()

# 整理输出
print(f"""
目标企业信息:
- Status: {status}
- Registration: {registration}
- State ID: {state_id}
- Business type: {business_type}
- Member: {member}
""")

这种方式更高效,适合处理大规模的网页内容。

内容的提问来源于stack exchange,提问作者Koj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 06:34:58