使用Selenium/Beautiful Soup爬取动态网站时丢失重复名称项的解决方法
问题描述
我正在结合Selenium与Beautiful Soup爬取目标网站,网站包含动态加载内容。目前存在部分数据未被爬取的问题:部分项名称重复(但链接内内容不同),导致输出中这些重复项被跳过丢失。需要将重复名称的项自动重命名为名称-1、名称-2的格式,后续还要爬取链接内容填充空字典。
以下是我的代码:
import json import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.edge.service import Service as EdgeService from webdriver_manager.microsoft import EdgeChromiumDriverManager from selenium.webdriver.edge.options import Options options = Options() options.add_argument("--start-maximized") URL = "https://atlantelavoro.inapp.org/atlante_professioni.php" driver = webdriver.Edge(service=EdgeService(EdgeChromiumDriverManager().install()), options=options) action_chains = ActionChains(driver) driver.get(URL) driver.implicitly_wait(0.5) elements = driver.find_elements(By.XPATH, "//a[@role='button']") for element in elements: action_chains.move_to_element(element).perform() element.click() ue = driver.find_elements(By.CSS_SELECTOR, "h4.panel-title4 a[role='button']") for element in ue: action_chains.move_to_element(element).perform() element.click() soup = BeautifulSoup(driver.page_source, "html.parser") out = {} tree = soup.find_all(class_="panel panel-info") for leaf in tree: out[leaf.select_one(".panel-title").get_text(strip=True)] = [e.get_text(strip=True) for e in leaf.select("h4.panel-title3 .ml-2")] if leaf.find_all("div", class_="panel-warning") != []: tutti = leaf.find_all("div", class_="panel-warning")[-1].select("div.panel-primary") ok = {} for t in tutti: ok[t.select_one("h4.panel-title4").get_text(strip=True)]= {e.get_text(strip=True): {} for e in t.select("h4.panel-title5")} for k, v in out.items(): if v == []: out[k] = ["Sezione in aggiornamento"] elif v[-1] == "Tutti": v[v.index("Tutti")] = {"Tutti": ok}
问题示例
以分类“01.\nAgricoltura, silvicoltura e pesca(11)”为例,括号内数字表示应包含11个项,但当前输出仅保留不重复的6个:
{'Addetto area della produzione': {}, 'Addetto conduzione macchine agricole': {}, 'Addetto conduzione macchine agricole (liv.2°)': {}, 'Addetto in allevamenti': {}, 'Addetto in aziende da latte e lattiero casearie': {}, 'Addetto in aziende orto-floro-frutticole': {}}
期望输出(重复项自动重命名):
{'Addetto area della produzione': {}, 'Addetto conduzione macchine agricole': {}, 'Addetto conduzione macchine agricole-1': {}, 'Addetto conduzione macchine agricole-2': {}, 'Addetto conduzione macchine agricole (liv.2°)': {}, 'Addetto in allevamenti': {}, 'Addetto in allevamenti-1': {}, 'Addetto in aziende da latte e lattiero casearie': {}, 'Addetto in aziende da latte e lattiero casearie-1': {}, 'Addetto in aziende orto-floro-frutticole': {}, 'Addetto in aziende orto-floro-frutticole-1': {}}
解决方案
问题根源是Python字典的键具有唯一性,直接用字典推导式会覆盖重复键。我们需要遍历元素时统计每个名称的出现次数,动态生成不重复的键:
修改核心代码片段
替换原代码中构建ok字典的循环逻辑:
# 原代码中这一行: # ok[t.select_one("h4.panel-title4").get_text(strip=True)]= {e.get_text(strip=True): {} for e in t.select("h4.panel-title5")} # 替换为以下代码: title4_text = t.select_one("h4.panel-title4").get_text(strip=True) job_dict = {} name_counter = {} # 遍历每个职业元素,处理重复名称 for e in t.select("h4.panel-title5"): job_name = e.get_text(strip=True) if job_name not in name_counter: name_counter[job_name] = 0 job_dict[job_name] = {} else: name_counter[job_name] += 1 new_job_name = f"{job_name}-{name_counter[job_name]}" job_dict[new_job_name] = {} ok[title4_text] = job_dict
完整修改后的代码
import json import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.edge.service import Service as EdgeService from webdriver_manager.microsoft import EdgeChromiumDriverManager from selenium.webdriver.edge.options import Options options = Options() options.add_argument("--start-maximized") URL = "https://atlantelavoro.inapp.org/atlante_professioni.php" driver = webdriver.Edge(service=EdgeService(EdgeChromiumDriverManager().install()), options=options) action_chains = ActionChains(driver) driver.get(URL) driver.implicitly_wait(0.5) elements = driver.find_elements(By.XPATH, "//a[@role='button']") for element in elements: action_chains.move_to_element(element).perform() element.click() ue = driver.find_elements(By.CSS_SELECTOR, "h4.panel-title4 a[role='button']") for element in ue: action_chains.move_to_element(element).perform() element.click() soup = BeautifulSoup(driver.page_source, "html.parser") out = {} tree = soup.find_all(class_="panel panel-info") for leaf in tree: out[leaf.select_one(".panel-title").get_text(strip=True)] = [e.get_text(strip=True) for e in leaf.select("h4.panel-title3 .ml-2")] if leaf.find_all("div", class_="panel-warning") != []: tutti = leaf.find_all("div", class_="panel-warning")[-1].select("div.panel-primary") ok = {} for t in tutti: title4_text = t.select_one("h4.panel-title4").get_text(strip=True) job_dict = {} name_counter = {} for e in t.select("h4.panel-title5"): job_name = e.get_text(strip=True) if job_name not in name_counter: name_counter[job_name] = 0 job_dict[job_name] = {} else: name_counter[job_name] += 1 new_job_name = f"{job_name}-{name_counter[job_name]}" job_dict[new_job_name] = {} ok[title4_text] = job_dict for k, v in out.items(): if v == []: out[k] = ["Sezione in aggiornamento"] elif v[-1] == "Tutti": v[v.index("Tutti")] = {"Tutti": ok}
后续爬取链接内容的扩展
如果需要填充空字典,可以在遍历e元素时获取链接并爬取内容:
# 在遍历e的循环中添加: job_link = e.get("href") # 用Selenium或requests请求链接解析内容,示例: # driver.get(job_link) # content = driver.page_source # 存入字典 job_dict[new_job_name if 'new_job_name' in locals() else job_name] = {"content": content}
内容的提问来源于stack exchange,提问作者sickboy83
相关产品推荐
相关产品推荐

