You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby中使用open-uri获取网页内容显示Loading...的解决方法

How to Fetch Dynamically Rendered HTML (Replace "Loading..." Content)

Hey there! The problem you're facing is super common when scraping modern websites: the content you want is dynamically loaded and rendered with JavaScript, but open-uri only grabs the raw initial HTML sent by the server—this is why you see the "Loading..." placeholder instead of the actual content. Browsers automatically execute the page's JavaScript to load and render the data, but open-uri doesn't do that.

Here are two reliable solutions to get the fully rendered HTML:

Solution 1: Use Selenium WebDriver (Simulates a Real Browser)

Selenium lets you control a real browser (like Chrome) programmatically, which will run all the page's JavaScript and render the complete content.

Step 1: Install the Gem

gem install selenium-webdriver

Step 2: Updated Code

require 'selenium-webdriver'

# Configure Headless Chrome (no visible browser window)
chrome_options = Selenium::WebDriver::Chrome::Options.new
chrome_options.add_argument('--headless=new')
chrome_options.add_argument('--disable-gpu') # Optional, for broader compatibility

# Initialize the browser driver
driver = Selenium::WebDriver.for :chrome, options: chrome_options

begin
  # Navigate to the target page
  driver.get('https://www.cargurus.com/Cars/l-Used-Mazda-MAZDASPEED6-d841')

  # Wait for dynamic content to load (replace the selector with your target element's CSS)
  wait = Selenium::WebDriver::Wait.new(timeout: 10)
  wait.until { driver.find_element(css: '.listing-row') } # Example selector for car listings

  # Grab the fully rendered HTML
  full_html = driver.page_source

  # Save to your file
  File.open("testing.txt", "w") do |file|
    file.puts full_html
  end
ensure
  # Always close the browser to free system resources
  driver.quit
end

Solution 2: Use Ferrum (Lightweight Headless Chrome Alternative)

Ferrum is a lighter Ruby library that interacts directly with Chrome's DevTools Protocol, so you don't need separate driver binaries.

Step 1: Install the Gem

gem install ferrum

Step 2: Updated Code

require 'ferrum'

# Initialize a Headless Chrome browser
browser = Ferrum::Browser.new(headless: true)

begin
  # Go to the target page
  browser.goto('https://www.cargurus.com/Cars/l-Used-Mazda-MAZDASPEED6-d841')

  # Wait for all network requests to finish (ensures content is fully loaded)
  browser.wait_for_network_idle(timeout: 10)

  # Get the fully rendered HTML body
  full_html = browser.body.html

  # Save to your file
  File.open("testing.txt", "w") do |file|
    file.puts full_html
  end
ensure
  # Close the browser session
  browser.quit
end

Quick Tips

  • Adjust Selectors: Replace the example CSS selector (.listing-row) with the actual selector for your target content—this makes sure the script waits until the data you care about is loaded.
  • Anti-Scraping Heads-Up: Some websites block automated browsers. You might need to add a custom User-Agent string or small delays between requests to avoid being blocked.
  • Dependencies: Both methods require Chrome (or Chromium) to be installed on your system.

内容的提问来源于stack exchange,提问作者user255580

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:54:04