技术需求:从多结果页面提取指定地址的目标企业数据
定向提取特定地址企业信息的解决方案
没问题,我来帮你搞定这个需求!针对你提供的HTML片段,我们只需要先定位到包含目标地址「311 South Swall Drive」的企业区块,再从中提取指定字段即可,具体实现思路和代码示例如下:
核心逻辑
每个企业的完整信息都包裹在<section class="org">标签中,我们可以:
- 遍历所有企业区块
- 检查区块内的
streetAddress文本是否匹配目标地址 - 找到匹配区块后,提取其中的Status、Registration等字段
实现方案(Python + BeautifulSoup)
这是最常用的网页解析方式,代码清晰易懂:
from bs4 import BeautifulSoup # 假设你的HTML源码存在这个变量里 html_content = """ <!-- 这里放你提供的HTML片段 --> """ soup = BeautifulSoup(html_content, 'html.parser') # 遍历所有企业区块 for org_section in soup.find_all('section', class_='org'): # 获取当前企业的街道地址 street_addr = org_section.find('span', itemprop='streetAddress') if street_addr and street_addr.text.strip() == '311 South Swall Drive': # 提取需要的字段 target_info = {} # 遍历所有属性行 for prop in org_section.find_all('p', class_='b-business-item_props'): title = prop.find('span', class_='b-business-item_title').text.strip().rstrip(':') value = prop.find('span', class_='b-business-item_value').text.strip() target_info[title] = value # 输出结果 print("目标企业信息:") for key, val in target_info.items(): print(f"- {key}: {val}") break # 找到目标后停止遍历
运行这段代码后,你会得到精准的结果:
- Status: Inactive
- Registration: Sep 26, 2006
- State ID: C2904860
- Business type: Articles of Incorporation
- Member: Ashwant Venkatram (President, inactive)
另一种方案:用XPath直接定位(适用于Scrapy、lxml)
如果你用Scrapy或者lxml这类支持XPath的工具,可以直接用一条XPath表达式定位到目标区块,再提取字段:
from lxml import etree tree = etree.HTML(html_content) # 定位目标企业区块 target_section = tree.xpath("//section[@class='org'][.//span[@itemprop='streetAddress' and text()='311 South Swall Drive']]")[0] # 提取字段 status = target_section.xpath(".//span[text()='Status:']/following-sibling::span/text()")[0].strip() registration = target_section.xpath(".//span[text()='Registration:']/following-sibling::span/text()")[0].strip() state_id = target_section.xpath(".//span[text()='State ID:']/following-sibling::span/text()")[0].strip() business_type = target_section.xpath(".//span[text()='Business type:']/following-sibling::span/text()")[0].strip() member = target_section.xpath(".//span[text()='Member:']/following-sibling::span/text()")[0].strip() # 整理输出 print(f""" 目标企业信息: - Status: {status} - Registration: {registration} - State ID: {state_id} - Business type: {business_type} - Member: {member} """)
这种方式更高效,适合处理大规模的网页内容。
内容的提问来源于stack exchange,提问作者Koj
相关产品推荐
相关产品推荐

