You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Nokogiri网页爬取:英超赛事信息爬取求助

Guide to Scraping Pinnacle Premier League Odds with Nokogiri

Hey there, I get it—scraping sites like Pinnacle can be way trickier than the simple static examples you find online. Let's break down the issues with your code and walk through a working approach.

First, Fix the Basic Parsing Issues

Your current code uses Nokogiri::XML.parse for an HTML page—this is a common mistake. HTML has different parsing rules than XML, so switch to Nokogiri::HTML.parse instead. Also, sites like Pinnacle often block requests that don't look like a real browser, and you might run into SSL errors. Let's adjust your code to handle those:

class BdcController < ApplicationController
  def bdc
    require 'nokogiri'
    require 'open-uri'
    require 'openssl'

    # Mimic a real browser user agent to avoid being blocked
    user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    
    begin
      # Fetch the page with proper headers and SSL handling
      doc = Nokogiri::HTML(open(
        'https://www.pinnacle.com/en/odds/match/soccer/england/england-premier-league',
        'User-Agent' => user_agent,
        ssl_verify_mode: OpenSSL::SSL::VERIFY_NONE # Note: This disables SSL checks (not ideal for production)
      ))

      # Now try to select elements—let's say we want match titles first
      match_titles = doc.css('.event-row__participants')
      match_titles.each do |title|
        puts title.text.strip
      end

    rescue OpenURI::HTTPError => e
      puts "HTTP Error: #{e.message}"
    rescue Exception => e
      puts "Error: #{e.message}"
    end
  end
end

Why This Might Still Not Work: Dynamic Content

Here's the big catch: Pinnacle loads most of its odds data dynamically using JavaScript. Nokogiri only parses the static HTML sent from the server—any content loaded after the initial page load (like the actual odds) won't be present in the doc object.

To get around this, you need to use a tool that can render the JavaScript and fully load the page before parsing it. The easiest way in Ruby is to use the ruby-puppeteer gem, which controls a headless Chrome browser to render the page.

Step-by-Step with Ruby-Puppeteer

  1. First, add the gem to your Gemfile:
gem 'ruby-puppeteer'

Then run bundle install.

  1. Update your controller code to use Puppeteer to render the page:
class BdcController < ApplicationController
  def bdc
    require 'nokogiri'
    require 'ruby_puppeteer'

    begin
      # Launch a headless Chrome browser
      RubyPuppeteer::Browser.launch(headless: true) do |browser|
        page = browser.new_page
        # Set a realistic user agent
        page.set_user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
        
        # Navigate to the page and wait for it to fully load
        page.goto('https://www.pinnacle.com/en/odds/match/soccer/england/england-premier-league', wait_until: 'networkidle2')
        
        # Get the fully rendered HTML
        html = page.content
        doc = Nokogiri::HTML(html)

        # Now extract the data you need—let's grab match details and odds
        doc.css('.event-row').each do |row|
          # Get match participants
          participants = row.css('.event-row__participants').text.strip
          # Get moneyline odds (adjust the selector based on current page structure)
          moneyline_odds = row.css('.odds-cell__odd').map(&:text).join(' | ')
          
          puts "#{participants} -> Odds: #{moneyline_odds}"
        end
      end

    rescue Exception => e
      puts "Error: #{e.message}"
    end
  end
end

Important Notes

  • Selector Changes: Pinnacle might update their HTML structure at any time, so you'll need to inspect the page (using Chrome DevTools) to find the latest CSS selectors for the data you want. Right now, .event-row is the container for each match, .event-row__participants is the team names, and .odds-cell__odd is the odds values—but these could change.
  • Rate Limiting: Don't spam requests to Pinnacle—they'll block your IP. Add delays between requests if you're scraping multiple pages.
  • Terms of Service: Always check a site's Terms of Service before scraping. Pinnacle might prohibit scraping their content, so make sure you're compliant.

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:29:58