You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Kotlin中使用Jsoup爬取新加坡博彩网站奖池金额失败求助

Fixing Your Kotlin + Jsoup Jackpot Scraper for Singapore Pools

Hey Jake, let's get that $1M estimated jackpot amount pulled correctly—your current code is tripping up on a selector syntax issue, plus a couple of easy fixes to avoid getting blocked by the site. Let's break this down:

First, the Core Issue: Your Selector Is Wrong

Your original selector div.slab__text slab__text--highlight uses a space between class names, which tells Jsoup to look for a child element with slab__text--highlight inside a div.slab__text. But if the jackpot element has both classes on the same div, you need to chain them with dots: div.slab__text.slab__text--highlight.

Fixed Code with Key Improvements

Here's a revised version that addresses selector syntax, adds browser emulation, and handles edge cases:

import org.jsoup.Jsoup

fun main() {
    try {
        // Mimic a real browser to avoid being blocked by the site's anti-scraping checks
        val doc = Jsoup.connect("https://online.singaporepools.com/lottery/en/home")
            .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
            .timeout(10000) // Add a timeout to avoid hanging on slow requests
            .get()

        // Target the exact element with the jackpot amount (adjust selector if page structure changes)
        val jackpotElement = doc.selectFirst("div.slab__text.slab__text--highlight")

        jackpotElement?.let {
            val cleanJackpot = it.text().trim()
            println("Estimated Jackpot: $cleanJackpot")
        } ?: run {
            println("Failed to find the jackpot element! Here's what to check:")
            println("1. Did the Singapore Pools page structure change?")
            println("2. Try printing the page's slab sections to debug:\n${doc.select("div.slab").take(3).joinToString("\n")}")
        }
    } catch (e: Exception) {
        println("Error scraping the site: ${e.message}")
        e.printStackTrace()
    }
}

Why This Works

  • Corrected Selector: div.slab__text.slab__text--highlight targets a div that has both slab__text and slab__text--highlight classes (the correct way to select elements with multiple classes).
  • User-Agent Header: Many sites block requests without a valid browser user agent—adding this makes your request look like it's coming from a real Chrome browser.
  • Error Handling: The selectFirst + let/?: combo prevents null pointer exceptions and gives you debug hints if the element can't be found.

If It Still Doesn't Work: Check for Dynamic Content

If the jackpot amount loads after the initial page (via JavaScript), Jsoup won't see it—since Jsoup only parses static HTML. In that case, you'll need a tool like Selenium to simulate a full browser that executes JavaScript:

import org.openqa.selenium.chrome.ChromeDriver
import org.openqa.selenium.By

fun main() {
    // Make sure you have ChromeDriver downloaded and its path set
    System.setProperty("webdriver.chrome.driver", "/path/to/your/chromedriver")
    
    val driver = ChromeDriver()
    try {
        driver.get("https://online.singaporepools.com/lottery/en/home")
        // Wait a second for JS to load the content (or use explicit waits for reliability)
        Thread.sleep(1000)
        
        val jackpotAmount = driver.findElement(By.cssSelector("div.slab__text.slab__text--highlight")).text
        println("Estimated Jackpot: $jackpotAmount")
    } finally {
        driver.quit() // Always close the browser when done
    }
}

Quick Note on Compliance

Before scraping, double-check Singapore Pools' terms of service—make sure your use case is allowed to avoid any legal issues.

内容的提问来源于stack exchange,提问作者jake wong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:29:48