You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium元素类型异常求助:Ubuntu环境下网页时刻表爬取失败

解决OASA时刻表爬取中提取单个时间的问题

看起来你遇到的核心问题是每个<li>元素包含多个发车时间,而你的代码只是把每个<li>作为一个整体元素,所以才会得到19个元素而非59个单独的时间。另外可能还存在页面加载不充分的问题,导致部分时间元素没被捕获到。我来帮你一步步解决:

问题分析

从你提供的HTML结构可以看到,每个<li>里有多个类似07:10、07:25的时间字符串,还有一个按钮的数字(比如07)。当前代码直接获取<li>列表,自然无法拆分出单个时间;另外隐式等待2秒可能不足以让页面完全渲染所有时刻表内容,导致实际能捕获的<li>数量不足。

解决方案

1. 改用显式等待确保页面加载完成

显式等待可以更精准地等待目标元素加载完毕,避免因页面动态渲染导致的元素遗漏。我们可以等待第一个时刻表<li>出现,或者等待所有相关元素加载完成。

2. 遍历每个<li>,提取并拆分时间字符串

对每个<li>元素,先获取其文本内容,然后通过字符串处理拆分出所有符合时间格式的内容(过滤掉按钮的数字和空字符)。

修改后的完整代码

import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.firefox.FirefoxDriver;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;

import java.time.Duration;
import java.util.ArrayList;
import java.util.List;
import java.util.regex.Pattern;

public class OASAScraper {
    public static void main(String[] args) {
        WebDriver driver = new FirefoxDriver();
        // 改用显式等待,设置10秒超时
        WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10));
        
        try {
            driver.get("http://telematics.oasa.gr/#main");
            
            // 显式等待下拉选项可点击,然后选择021线路
            WebElement lineOption = wait.until(ExpectedConditions.elementToBeClickable(By.xpath("//option[contains(.,'021')]")));
            lineOption.click();
            
            // 显式等待时刻表列表加载完成
            wait.until(ExpectedConditions.presenceOfElementLocated(By.xpath("//div/ul/li[@class='list-group-item scheduleEntryL']")));
            
            // 获取所有时刻表的li元素
            List<WebElement> timeRows = driver.findElements(By.xpath("//div/ul/li[@class='list-group-item scheduleEntryL']"));
            
            // 用于存储所有单个时间
            List<String> allTimes = new ArrayList<>();
            // 正则匹配时间格式(HH:MM)
            Pattern timePattern = Pattern.compile("\\d{2}:\\d{2}");
            
            for (WebElement row : timeRows) {
                String rowText = row.getText();
                // 按空格分割文本,然后过滤出符合时间格式的内容
                String[] parts = rowText.split("\\s+");
                for (String part : parts) {
                    if (timePattern.matcher(part).matches()) {
                        allTimes.add(part);
                    }
                }
            }
            
            // 输出结果
            System.out.println("总时间数量: " + allTimes.size());
            System.out.println("所有时间列表:");
            for (String time : allTimes) {
                System.out.println(time);
            }
            
            // 保存到.txt文件(可选,需导入java.nio.file相关包)
            // Files.write(Paths.get("oasa_times.txt"), allTimes, StandardCharsets.UTF_8);
            
        } catch (Exception e) {
            e.printStackTrace();
        } finally {
            driver.quit();
        }
    }
}

代码说明

  • 显式等待:使用WebDriverWait等待元素可点击或存在,确保页面动态渲染完成,避免遗漏元素。
  • 正则匹配时间:用\\d{2}:\\d{2}的正则表达式精准匹配HH:MM格式的时间,过滤掉按钮的数字(如07)和空字符。
  • 字符串拆分:通过split("\\s+")按任意数量的空格分割文本,再逐个检查每个部分是否符合时间格式。
  • 资源清理:在finally块中关闭浏览器,避免资源泄漏。

这样处理后,你应该能得到完整的59个发车时间,并且每个时间都是单独的字符串,方便保存到.txt文件或进一步处理。

内容的提问来源于stack exchange,提问作者YourHelper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 06:57:32