如何在ggplot中将因子转数值并绘制日期-企业数量统计图?
Hey there! Let's work through your problems step by step—first fixing that messy ggplot chart, then tackling the factor-to-numeric conversion, and even cleaning up your web scraping code a bit.
1. Fixing the ggplot Bar Chart (Counting Firms Per Date)
The issue with your current plot is that you're passing the raw Firm name column directly to the y-axis. ggplot treats each unique firm name as a separate category, which is why your y-axis is cluttered with every business name instead of showing counts.
First, you need to summarize your data to count how many firms exist per date, then plot that summary:
# Step 1: Calculate number of firms per date using dplyr nordland_counts <- nordland %>% count(Date, name = "Number_of_firms") # Name the count column explicitly # Step 2: Build the corrected plot plot <- ggplot(nordland_counts, aes(x = Date, y = Number_of_firms)) + geom_col(fill = "#4C72B0") # Add a clean color (optional) labs( x = "Date", y = "Number of firms", title = "Number of new firms per month" ) + theme_minimal() # Optional: Makes the chart look cleaner
A quick note: In ggplot, you don't need to use nordland$Date in the aes() call—just use the column name directly, since you've already specified the dataset with data = nordland_counts.
2. Converting Factors to Numeric in R
When converting factors to numeric, you have to be careful: using as.numeric() directly on a factor returns the index of the factor's level, not the actual numeric value stored in the factor. Here's the correct way:
# For a standalone factor column numeric_column <- as.numeric(as.character(your_factor_column)) # In a tidyverse pipeline (if you're modifying a data frame) your_data <- your_data %>% mutate(numeric_column = as.numeric(as.character(your_factor_column)))
Also, looking at your scraping code—you converted Firmname to a factor, but this isn't necessary for firm names (it can actually cause headaches later). You can safely remove the Firmname <- as.factor(Firmname) line.
3. Cleaning Up Your Web Scraping Code
Your scraping code works, but we can make it more concise and readable using tidyverse functions:
library(rvest) library(tidyverse) url <- "https://w2.brreg.no/kunngjoring/kombisok.jsp?datoFra=01.01.2019&datoTil=25.09.2019&id_region=100&id_fylke=-+-+-&id_niva1=2&id_bransje1=0" nordland <- read_html(url) %>% # Extract and clean firm names and dates in one go tibble( `Firm name` = html_nodes(., "td td:nth-child(2) p") %>% html_text() %>% str_remove_all("\n| ") %>% # Remove newlines and spaces in one step .[-1], # Drop the first row (header) Date = html_nodes(., "td:nth-child(6) p") %>% html_text() %>% as.Date("%d.%m.%Y") %>% .[-1] # Drop the first row to match firm names ) %>% slice(1:1052) # Keep the first 1052 rows as you did originally
This avoids creating multiple intermediate data frames and uses str_remove_all() instead of repeated gsub() calls for cleaner code.
内容的提问来源于stack exchange,提问作者Kristian Bjerke

