You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取需求:为Allocine网站爬取结果添加数据分隔符

Fixing Your Allocine Web Scraping Script

Hey Pierre, let's get your Allocine scraping script producing the clean, structured output you want! The main problem with your current code is that you’re adding each piece of data (movie name, ratings) as separate items to your array—so your data ends up ungrouped, which is why your output looks jumbled. Here’s how to fix it:

Key Issues in Your Original Code

  • You’re pushing movie names and ratings as individual elements into array, resulting in a flat list like ["Coco", "4,14,6", "Forrest Gump", "2,64,6"] instead of grouped entries.
  • When writing to CSV, you’re inserting each element as a new row, which breaks the movie-rating association.

Corrected Code

require 'open-uri'
require 'nokogiri'
require 'csv'

# Initialize an array to hold grouped movie data (each entry is [name, rating1, rating2])
movies = []

# Iterate over pages 1 to 10 (using Ruby's preferred each loop instead of for)
(1..10).each do |page_num|
  url = "http://www.allocine.fr/film/meilleurs//?page=#{page_num}"
  html_file = open(url).read
  html_doc = Nokogiri::HTML(html_file)

  # Process each movie element on the page
  html_doc.search('.img_side_content').each do |movie_element|
    # Extract and clean the movie name
    movie_name = movie_element.search('.no_underline').inner_text.strip

    # Extract both ratings (assuming .note targets the two rating elements on the page)
    ratings = movie_element.search('.note').map { |note| note.inner_text.strip }

    # Add to movies array: use ratings if available, else default to 'N/A'
    movie_entry = if ratings.size >= 2
                    [movie_name, ratings[0], ratings[1]]
                  else
                    [movie_name, 'N/A', 'N/A']
                  end
    movies << movie_entry
  end
end

# Print the structured output to console (matches your desired format)
puts "Movie Name Rating 1 Rating 2"
movies.each do |movie|
  puts "#{movie[0]} #{movie[1]} #{movie[2]}"
end

# Write the structured data to CSV
csv_options = { col_sep: ',', force_quotes: true, quote_char: '"' }
filepath = 'allocine.csv'

CSV.open(filepath, 'wb', csv_options) do |csv|
  # Optional: Add a header row to make the CSV more readable
  csv << ['Movie Name', 'Rating 1', 'Rating 2']
  # Write each movie's data as a single row
  movies.each { |movie| csv << movie }
end

What Changed?

  • Grouped data structure: Each movie’s details are stored as a sub-array, ensuring name and ratings stay linked.
  • Ruby-idiomatic loops: Replaced for i in 1..10 with (1..10).each for cleaner, more readable code.
  • Whitespace cleanup: .strip() removes extra spaces/newlines from scraped text to avoid messy output.
  • Robust rating handling: Added a fallback for movies missing ratings so your script doesn’t break unexpectedly.
  • Proper CSV writing: Each sub-array is written as a single row, with an optional header for clarity.

内容的提问来源于stack exchange,提问作者Pierre Laburthe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:43:26