使用Nokogiri网页爬取:英超赛事信息爬取求助
Hey there, I get it—scraping sites like Pinnacle can be way trickier than the simple static examples you find online. Let's break down the issues with your code and walk through a working approach.
First, Fix the Basic Parsing Issues
Your current code uses Nokogiri::XML.parse for an HTML page—this is a common mistake. HTML has different parsing rules than XML, so switch to Nokogiri::HTML.parse instead. Also, sites like Pinnacle often block requests that don't look like a real browser, and you might run into SSL errors. Let's adjust your code to handle those:
class BdcController < ApplicationController def bdc require 'nokogiri' require 'open-uri' require 'openssl' # Mimic a real browser user agent to avoid being blocked user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" begin # Fetch the page with proper headers and SSL handling doc = Nokogiri::HTML(open( 'https://www.pinnacle.com/en/odds/match/soccer/england/england-premier-league', 'User-Agent' => user_agent, ssl_verify_mode: OpenSSL::SSL::VERIFY_NONE # Note: This disables SSL checks (not ideal for production) )) # Now try to select elements—let's say we want match titles first match_titles = doc.css('.event-row__participants') match_titles.each do |title| puts title.text.strip end rescue OpenURI::HTTPError => e puts "HTTP Error: #{e.message}" rescue Exception => e puts "Error: #{e.message}" end end end
Why This Might Still Not Work: Dynamic Content
Here's the big catch: Pinnacle loads most of its odds data dynamically using JavaScript. Nokogiri only parses the static HTML sent from the server—any content loaded after the initial page load (like the actual odds) won't be present in the doc object.
To get around this, you need to use a tool that can render the JavaScript and fully load the page before parsing it. The easiest way in Ruby is to use the ruby-puppeteer gem, which controls a headless Chrome browser to render the page.
Step-by-Step with Ruby-Puppeteer
- First, add the gem to your
Gemfile:
gem 'ruby-puppeteer'
Then run bundle install.
- Update your controller code to use Puppeteer to render the page:
class BdcController < ApplicationController def bdc require 'nokogiri' require 'ruby_puppeteer' begin # Launch a headless Chrome browser RubyPuppeteer::Browser.launch(headless: true) do |browser| page = browser.new_page # Set a realistic user agent page.set_user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") # Navigate to the page and wait for it to fully load page.goto('https://www.pinnacle.com/en/odds/match/soccer/england/england-premier-league', wait_until: 'networkidle2') # Get the fully rendered HTML html = page.content doc = Nokogiri::HTML(html) # Now extract the data you need—let's grab match details and odds doc.css('.event-row').each do |row| # Get match participants participants = row.css('.event-row__participants').text.strip # Get moneyline odds (adjust the selector based on current page structure) moneyline_odds = row.css('.odds-cell__odd').map(&:text).join(' | ') puts "#{participants} -> Odds: #{moneyline_odds}" end end rescue Exception => e puts "Error: #{e.message}" end end end
Important Notes
- Selector Changes: Pinnacle might update their HTML structure at any time, so you'll need to inspect the page (using Chrome DevTools) to find the latest CSS selectors for the data you want. Right now,
.event-rowis the container for each match,.event-row__participantsis the team names, and.odds-cell__oddis the odds values—but these could change. - Rate Limiting: Don't spam requests to Pinnacle—they'll block your IP. Add delays between requests if you're scraping multiple pages.
- Terms of Service: Always check a site's Terms of Service before scraping. Pinnacle might prohibit scraping their content, so make sure you're compliant.
内容的提问来源于stack exchange,提问作者John

