You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python 2.7+Selenium爬取HTML:如何匹配公司下的人员数据

解决不定数量Person归属Company的爬取问题

嘿,这个场景我太熟悉了!很多网站的列表都会这么设计——公司行后面跟着数量不定的联系人行,咱们用Selenium结合Python 2.7完全可以精准搞定,下面给你两种靠谱的思路和代码示例:

方法一:按顺序遍历分组(最直观不易出错)

这种方法是先把所有的<tr>元素都取出来,然后按顺序遍历:遇到company类的行就开启一个新分组,把后续的person行都归到这个分组里,直到遇到下一个company行为止。完全不用纠结数量,自动适配1-4个甚至更多的情况。

代码示例:

from selenium import webdriver
import csv

# 初始化浏览器(这里用Firefox,你可以换成Chrome等)
driver = webdriver.Firefox()
driver.get("你的目标网站URL")

# 获取所有<tr>元素
all_rows = driver.find_elements_by_tag_name("tr")

# 初始化数据存储结构
company_groups = []
current_company = None
current_persons = []

# 遍历所有行进行分组
for row in all_rows:
    row_class = row.get_attribute("class")
    if row_class == "company":
        # 如果之前有未保存的公司和人员,先存入列表
        if current_company is not None:
            company_groups.append({
                "company_name": current_company.text.strip(),
                "persons": [p.text.strip() for p in current_persons]
            })
        # 切换到新的公司分组
        current_company = row
        current_persons = []
    elif row_class == "person":
        # 确保当前人员属于某个公司(避免页面开头就是person的异常情况)
        if current_company is not None:
            current_persons.append(row)

# 处理最后一组公司和人员(避免循环结束后遗漏)
if current_company is not None:
    company_groups.append({
        "company_name": current_company.text.strip(),
        "persons": [p.text.strip() for p in current_persons]
    })

# 将数据写入CSV文件(Python 2.7注意用"wb"模式,避免多余空行)
with open("company_persons.csv", "wb") as csv_file:
    writer = csv.writer(csv_file)
    # 写入表头
    writer.writerow(["Company", "Person"])
    # 按"公司-人员"的对应关系逐行写入
    for group in company_groups:
        comp_name = group["company_name"]
        for person in group["persons"]:
            writer.writerow([comp_name, person])

# 关闭浏览器
driver.quit()

方法二:用XPath精准定位同组Person

如果你熟悉XPath语法,也可以直接针对每个company元素,定位它后面所有属于同一组的person元素——也就是当前company之后、下一个company之前的所有person行,逻辑更紧凑。

核心代码片段:

# 获取所有company行
companies = driver.find_elements_by_css_selector("tr.company")

company_groups = []
for comp in companies:
    comp_name = comp.text.strip()
    # 用XPath筛选当前company之后的同组person
    persons = comp.find_elements_by_xpath(
        "./following-sibling::tr[@class='person'][count(preceding-sibling::tr[@class='company']) = count(current()/preceding-sibling::tr[@class='company']) + 1]"
    )
    person_list = [p.text.strip() for p in persons]
    company_groups.append({"company_name": comp_name, "persons": person_list})

# 后续写入CSV的逻辑和方法一一致

额外提示:

  • 如果你需要提取的不是整行文本,而是行内的具体字段(比如公司全称、人员姓名、联系方式),只需要把row.text.strip()换成对应的元素定位即可,比如row.find_element_by_css_selector("td.company-name").text.strip()。
  • 页面加载慢的话,可以加入WebDriverWait等待元素加载完成,避免出现找不到元素的错误。

内容的提问来源于stack exchange,提问作者Rashid Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:58:03