You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium/Beautiful Soup爬取动态网站时丢失重复名称项的解决方法

问题描述

我正在结合Selenium与Beautiful Soup爬取目标网站,网站包含动态加载内容。目前存在部分数据未被爬取的问题:部分项名称重复(但链接内内容不同),导致输出中这些重复项被跳过丢失。需要将重复名称的项自动重命名为名称-1、名称-2的格式,后续还要爬取链接内容填充空字典。

以下是我的代码:

import json
import requests
from bs4 import BeautifulSoup

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.edge.service import Service as EdgeService
from webdriver_manager.microsoft import EdgeChromiumDriverManager
from selenium.webdriver.edge.options import Options

options = Options()
options.add_argument("--start-maximized")

URL = "https://atlantelavoro.inapp.org/atlante_professioni.php"

driver = webdriver.Edge(service=EdgeService(EdgeChromiumDriverManager().install()), options=options)
action_chains = ActionChains(driver)
driver.get(URL)
driver.implicitly_wait(0.5)
elements = driver.find_elements(By.XPATH, "//a[@role='button']")
for element in elements:
    action_chains.move_to_element(element).perform()
    element.click()

ue = driver.find_elements(By.CSS_SELECTOR, "h4.panel-title4 a[role='button']")

for element in ue:
    action_chains.move_to_element(element).perform()
    element.click()


soup = BeautifulSoup(driver.page_source, "html.parser")
out = {}
tree = soup.find_all(class_="panel panel-info")

for leaf in tree:
    out[leaf.select_one(".panel-title").get_text(strip=True)] = [e.get_text(strip=True) for e in leaf.select("h4.panel-title3 .ml-2")]
    if leaf.find_all("div", class_="panel-warning") != []:
        tutti = leaf.find_all("div", class_="panel-warning")[-1].select("div.panel-primary")
        ok = {}
        for t in tutti:
            ok[t.select_one("h4.panel-title4").get_text(strip=True)]= {e.get_text(strip=True): {} for e in t.select("h4.panel-title5")}

    for k, v in out.items():
        if v == []:
            out[k] = ["Sezione in aggiornamento"]
        elif v[-1] == "Tutti":
            v[v.index("Tutti")] = {"Tutti": ok}

问题示例

以分类“01.\nAgricoltura, silvicoltura e pesca(11)”为例,括号内数字表示应包含11个项,但当前输出仅保留不重复的6个:

{'Addetto area della produzione': {}, 'Addetto conduzione macchine agricole': {}, 'Addetto conduzione macchine agricole (liv.2°)': {}, 'Addetto in allevamenti': {}, 'Addetto in aziende da latte e lattiero casearie': {}, 'Addetto in aziende orto-floro-frutticole': {}}

期望输出(重复项自动重命名):

{'Addetto area della produzione': {}, 'Addetto conduzione macchine agricole': {}, 'Addetto conduzione macchine agricole-1': {}, 'Addetto conduzione macchine agricole-2': {}, 'Addetto conduzione macchine agricole (liv.2°)': {}, 'Addetto in allevamenti': {}, 'Addetto in allevamenti-1': {}, 'Addetto in aziende da latte e lattiero casearie': {}, 'Addetto in aziende da latte e lattiero casearie-1': {}, 'Addetto in aziende orto-floro-frutticole': {}, 'Addetto in aziende orto-floro-frutticole-1': {}}

解决方案

问题根源是Python字典的键具有唯一性,直接用字典推导式会覆盖重复键。我们需要遍历元素时统计每个名称的出现次数,动态生成不重复的键:

修改核心代码片段

替换原代码中构建ok字典的循环逻辑:

# 原代码中这一行:
# ok[t.select_one("h4.panel-title4").get_text(strip=True)]= {e.get_text(strip=True): {} for e in t.select("h4.panel-title5")}

# 替换为以下代码:
title4_text = t.select_one("h4.panel-title4").get_text(strip=True)
job_dict = {}
name_counter = {}

# 遍历每个职业元素,处理重复名称
for e in t.select("h4.panel-title5"):
    job_name = e.get_text(strip=True)
    if job_name not in name_counter:
        name_counter[job_name] = 0
        job_dict[job_name] = {}
    else:
        name_counter[job_name] += 1
        new_job_name = f"{job_name}-{name_counter[job_name]}"
        job_dict[new_job_name] = {}

ok[title4_text] = job_dict

完整修改后的代码

import json
import requests
from bs4 import BeautifulSoup

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.edge.service import Service as EdgeService
from webdriver_manager.microsoft import EdgeChromiumDriverManager
from selenium.webdriver.edge.options import Options

options = Options()
options.add_argument("--start-maximized")

URL = "https://atlantelavoro.inapp.org/atlante_professioni.php"

driver = webdriver.Edge(service=EdgeService(EdgeChromiumDriverManager().install()), options=options)
action_chains = ActionChains(driver)
driver.get(URL)
driver.implicitly_wait(0.5)
elements = driver.find_elements(By.XPATH, "//a[@role='button']")
for element in elements:
    action_chains.move_to_element(element).perform()
    element.click()

ue = driver.find_elements(By.CSS_SELECTOR, "h4.panel-title4 a[role='button']")

for element in ue:
    action_chains.move_to_element(element).perform()
    element.click()


soup = BeautifulSoup(driver.page_source, "html.parser")
out = {}
tree = soup.find_all(class_="panel panel-info")

for leaf in tree:
    out[leaf.select_one(".panel-title").get_text(strip=True)] = [e.get_text(strip=True) for e in leaf.select("h4.panel-title3 .ml-2")]
    if leaf.find_all("div", class_="panel-warning") != []:
        tutti = leaf.find_all("div", class_="panel-warning")[-1].select("div.panel-primary")
        ok = {}
        for t in tutti:
            title4_text = t.select_one("h4.panel-title4").get_text(strip=True)
            job_dict = {}
            name_counter = {}
            for e in t.select("h4.panel-title5"):
                job_name = e.get_text(strip=True)
                if job_name not in name_counter:
                    name_counter[job_name] = 0
                    job_dict[job_name] = {}
                else:
                    name_counter[job_name] += 1
                    new_job_name = f"{job_name}-{name_counter[job_name]}"
                    job_dict[new_job_name] = {}
            ok[title4_text] = job_dict

    for k, v in out.items():
        if v == []:
            out[k] = ["Sezione in aggiornamento"]
        elif v[-1] == "Tutti":
            v[v.index("Tutti")] = {"Tutti": ok}

后续爬取链接内容的扩展

如果需要填充空字典,可以在遍历e元素时获取链接并爬取内容:

# 在遍历e的循环中添加:
job_link = e.get("href")
# 用Selenium或requests请求链接解析内容,示例:
# driver.get(job_link)
# content = driver.page_source
# 存入字典
job_dict[new_job_name if 'new_job_name' in locals() else job_name] = {"content": content}

内容的提问来源于stack exchange,提问作者sickboy83

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 06:37:39