使用Python Selenium提取酒店数据遇重复问题及HTML模板填充需求
问题1:修复内容重复与仅提取首个h3的问题
你遇到的问题大概率是遍历范围错误或误用单元素定位方法导致的,以下是修正后的Selenium提取逻辑:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("https://es.hoteles.com/ho227810/secrets-lanzarote-resort-spa-adults-only-18-yaiza-espana/") # 等待所有特征区块加载完成(根据页面实际结构调整选择器) wait = WebDriverWait(driver, 10) feature_blocks = wait.until(EC.presence_of_all_elements_located( (By.CSS_SELECTOR, "div.feature-group") # 假设每个h3和对应列表都包裹在该类的div中 )) hotel_features = [] for block in feature_blocks: # 在当前区块内提取h3文本 h3_text = block.find_element(By.TAG_NAME, "h3").text.strip() # 在当前区块内提取所有列表项 list_items = [item.text.strip() for item in block.find_elements(By.TAG_NAME, "li")] hotel_features.append({ "title": h3_text, "items": list_items }) driver.quit()
关键修正点:
- 用
find_elements(复数形式)定位所有特征区块,避免只获取单个元素 - 遍历每个区块时,在当前区块内部定位h3和列表项,防止跨区块重复提取
- 加入显式等待确保元素加载完成,避免因未加载导致的空数据或定位失败
问题2:自动填充HTML表格模板
可以通过字符串格式化或轻量模板引擎实现,以下是两种实用方案:
方案1:原生字符串格式化(适合简单模板)
假设你的HTML模板结构如下:
<table> <thead> <tr> <th>序号</th> <th>标题</th> <th>占位</th> <th>内容</th> </tr> </thead> <tbody> <!-- 动态行占位 --> </tbody> </table>
填充代码:
# 读取模板文件(或直接定义模板字符串) with open("template.html", "r", encoding="utf-8") as f: template = f.read() # 生成动态表格行 table_rows = "" for idx, feature in enumerate(hotel_features, 1): items_str = ", ".join(feature["items"]) table_rows += f""" <tr> <td>{idx}</td> <td>{feature["title"]}</td> <td></td> <td>{items_str}</td> </tr> """ # 替换模板占位并保存结果 final_html = template.replace("<!-- 动态行占位 -->", table_rows) with open("filled_hotel_table.html", "w", encoding="utf-8") as f: f.write(final_html)
方案2:Jinja2模板引擎(适合复杂模板)
- 安装依赖:
pip install jinja2 - 定义模板文件
table_template.html:
<table> <thead> <tr> <th>序号</th> <th>标题</th> <th>占位</th> <th>内容</th> </tr> </thead> <tbody> {% for idx, feature in enumerate(hotel_features, 1) %} <tr> <td>{{ idx }}</td> <td>{{ feature.title }}</td> <td></td> <td>{{ feature.items | join(', ') }}</td> </tr> {% endfor %} </tbody> </table>
- 渲染模板代码:
from jinja2 import Environment, FileSystemLoader # 加载模板环境 env = Environment(loader=FileSystemLoader(".")) template = env.get_template("table_template.html") # 渲染并保存结果 final_html = template.render(hotel_features=hotel_features) with open("filled_hotel_table.html", "w", encoding="utf-8") as f: f.write(final_html)
内容的提问来源于stack exchange,提问作者Felipe jimenez prieto
相关产品推荐
相关产品推荐

