Ruby中使用open-uri获取网页内容显示Loading...的解决方法
Hey there! The problem you're facing is super common when scraping modern websites: the content you want is dynamically loaded and rendered with JavaScript, but open-uri only grabs the raw initial HTML sent by the server—this is why you see the "Loading..." placeholder instead of the actual content. Browsers automatically execute the page's JavaScript to load and render the data, but open-uri doesn't do that.
Here are two reliable solutions to get the fully rendered HTML:
Solution 1: Use Selenium WebDriver (Simulates a Real Browser)
Selenium lets you control a real browser (like Chrome) programmatically, which will run all the page's JavaScript and render the complete content.
Step 1: Install the Gem
gem install selenium-webdriver
Step 2: Updated Code
require 'selenium-webdriver' # Configure Headless Chrome (no visible browser window) chrome_options = Selenium::WebDriver::Chrome::Options.new chrome_options.add_argument('--headless=new') chrome_options.add_argument('--disable-gpu') # Optional, for broader compatibility # Initialize the browser driver driver = Selenium::WebDriver.for :chrome, options: chrome_options begin # Navigate to the target page driver.get('https://www.cargurus.com/Cars/l-Used-Mazda-MAZDASPEED6-d841') # Wait for dynamic content to load (replace the selector with your target element's CSS) wait = Selenium::WebDriver::Wait.new(timeout: 10) wait.until { driver.find_element(css: '.listing-row') } # Example selector for car listings # Grab the fully rendered HTML full_html = driver.page_source # Save to your file File.open("testing.txt", "w") do |file| file.puts full_html end ensure # Always close the browser to free system resources driver.quit end
Solution 2: Use Ferrum (Lightweight Headless Chrome Alternative)
Ferrum is a lighter Ruby library that interacts directly with Chrome's DevTools Protocol, so you don't need separate driver binaries.
Step 1: Install the Gem
gem install ferrum
Step 2: Updated Code
require 'ferrum' # Initialize a Headless Chrome browser browser = Ferrum::Browser.new(headless: true) begin # Go to the target page browser.goto('https://www.cargurus.com/Cars/l-Used-Mazda-MAZDASPEED6-d841') # Wait for all network requests to finish (ensures content is fully loaded) browser.wait_for_network_idle(timeout: 10) # Get the fully rendered HTML body full_html = browser.body.html # Save to your file File.open("testing.txt", "w") do |file| file.puts full_html end ensure # Close the browser session browser.quit end
Quick Tips
- Adjust Selectors: Replace the example CSS selector (
.listing-row) with the actual selector for your target content—this makes sure the script waits until the data you care about is loaded. - Anti-Scraping Heads-Up: Some websites block automated browsers. You might need to add a custom User-Agent string or small delays between requests to avoid being blocked.
- Dependencies: Both methods require Chrome (or Chromium) to be installed on your system.
内容的提问来源于stack exchange,提问作者user255580

