使用Python 2.7+Selenium爬取HTML:如何匹配公司下的人员数据
解决不定数量Person归属Company的爬取问题
嘿,这个场景我太熟悉了!很多网站的列表都会这么设计——公司行后面跟着数量不定的联系人行,咱们用Selenium结合Python 2.7完全可以精准搞定,下面给你两种靠谱的思路和代码示例:
方法一:按顺序遍历分组(最直观不易出错)
这种方法是先把所有的<tr>元素都取出来,然后按顺序遍历:遇到company类的行就开启一个新分组,把后续的person行都归到这个分组里,直到遇到下一个company行为止。完全不用纠结数量,自动适配1-4个甚至更多的情况。
代码示例:
from selenium import webdriver import csv # 初始化浏览器(这里用Firefox,你可以换成Chrome等) driver = webdriver.Firefox() driver.get("你的目标网站URL") # 获取所有<tr>元素 all_rows = driver.find_elements_by_tag_name("tr") # 初始化数据存储结构 company_groups = [] current_company = None current_persons = [] # 遍历所有行进行分组 for row in all_rows: row_class = row.get_attribute("class") if row_class == "company": # 如果之前有未保存的公司和人员,先存入列表 if current_company is not None: company_groups.append({ "company_name": current_company.text.strip(), "persons": [p.text.strip() for p in current_persons] }) # 切换到新的公司分组 current_company = row current_persons = [] elif row_class == "person": # 确保当前人员属于某个公司(避免页面开头就是person的异常情况) if current_company is not None: current_persons.append(row) # 处理最后一组公司和人员(避免循环结束后遗漏) if current_company is not None: company_groups.append({ "company_name": current_company.text.strip(), "persons": [p.text.strip() for p in current_persons] }) # 将数据写入CSV文件(Python 2.7注意用"wb"模式,避免多余空行) with open("company_persons.csv", "wb") as csv_file: writer = csv.writer(csv_file) # 写入表头 writer.writerow(["Company", "Person"]) # 按"公司-人员"的对应关系逐行写入 for group in company_groups: comp_name = group["company_name"] for person in group["persons"]: writer.writerow([comp_name, person]) # 关闭浏览器 driver.quit()
方法二:用XPath精准定位同组Person
如果你熟悉XPath语法,也可以直接针对每个company元素,定位它后面所有属于同一组的person元素——也就是当前company之后、下一个company之前的所有person行,逻辑更紧凑。
核心代码片段:
# 获取所有company行 companies = driver.find_elements_by_css_selector("tr.company") company_groups = [] for comp in companies: comp_name = comp.text.strip() # 用XPath筛选当前company之后的同组person persons = comp.find_elements_by_xpath( "./following-sibling::tr[@class='person'][count(preceding-sibling::tr[@class='company']) = count(current()/preceding-sibling::tr[@class='company']) + 1]" ) person_list = [p.text.strip() for p in persons] company_groups.append({"company_name": comp_name, "persons": person_list}) # 后续写入CSV的逻辑和方法一一致
额外提示:
- 如果你需要提取的不是整行文本,而是行内的具体字段(比如公司全称、人员姓名、联系方式),只需要把
row.text.strip()换成对应的元素定位即可,比如row.find_element_by_css_selector("td.company-name").text.strip()。 - 页面加载慢的话,可以加入
WebDriverWait等待元素加载完成,避免出现找不到元素的错误。
内容的提问来源于stack exchange,提问作者Rashid Aziz
相关产品推荐
相关产品推荐

