R网页抓取后的数据表格格式化问题求助
Hey there! Let's work through this together. I spot a couple of small hiccups in your code that are throwing things off, plus we need to properly parse the CSV data you're grabbing from the page. Here's how to get your 4-column table sorted:
First, you have a syntax error in your variable name: page sources should be page_sources (spaces aren't allowed in R variable names). That's likely causing an error right out the gate.
Next, the content you're pulling from the page is a raw CSV string—just wrapping it in as.data.frame() won't split it into structured columns correctly. We need to use a CSV parser to turn that string into a proper data frame.
Here's the corrected code that first gets your three valid columns, then adds the empty column to reach your 4-column goal:
library(rvest) # The methods package isn't necessary here unless you're using specific methods later # Fetch the page and extract the raw CSV text page <- read_html("https://www.galmarley.com/prices/CSV/AUX/USD/600/Full") page_sources <- page %>% html_text() # Parse the CSV string into a structured data frame raw_data <- read.csv(text = page_sources) # Add an empty column to create your 4-column table final_data <- cbind(Empty_Column = "", raw_data) # Preview the result head(final_data)
A quick breakdown of what's happening:
read.csv(text = page_sources)takes the raw CSV string and converts it into a data frame with the correct columns from the source.cbind(Empty_Column = "", raw_data)adds a new empty column at the start (swap the order tocbind(raw_data, Empty_Column = "")if you want the empty column at the end instead).
If you prefer using tidyverse tools (since you're already using pipes with rvest), here's an alternative version:
library(rvest) library(dplyr) library(readr) final_data <- read_html("https://www.galmarley.com/prices/CSV/AUX/USD/600/Full") %>% html_text() %>% read_csv() %>% mutate(Empty_Column = "") # Adds empty column; use relocate() to adjust its position head(final_data)
This should give you exactly the table you need: one empty column plus the three valid columns from your scraped data.
内容的提问来源于stack exchange,提问作者adam.888

