You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在ggplot中将因子转数值并绘制日期-企业数量统计图?

Hey there! Let's work through your problems step by step—first fixing that messy ggplot chart, then tackling the factor-to-numeric conversion, and even cleaning up your web scraping code a bit.

1. Fixing the ggplot Bar Chart (Counting Firms Per Date)

The issue with your current plot is that you're passing the raw Firm name column directly to the y-axis. ggplot treats each unique firm name as a separate category, which is why your y-axis is cluttered with every business name instead of showing counts.

First, you need to summarize your data to count how many firms exist per date, then plot that summary:

# Step 1: Calculate number of firms per date using dplyr
nordland_counts <- nordland %>%
  count(Date, name = "Number_of_firms") # Name the count column explicitly

# Step 2: Build the corrected plot
plot <- ggplot(nordland_counts, aes(x = Date, y = Number_of_firms)) +
  geom_col(fill = "#4C72B0") # Add a clean color (optional)
  labs(
    x = "Date", 
    y = "Number of firms", 
    title = "Number of new firms per month"
  ) +
  theme_minimal() # Optional: Makes the chart look cleaner

A quick note: In ggplot, you don't need to use nordland$Date in the aes() call—just use the column name directly, since you've already specified the dataset with data = nordland_counts.

2. Converting Factors to Numeric in R

When converting factors to numeric, you have to be careful: using as.numeric() directly on a factor returns the index of the factor's level, not the actual numeric value stored in the factor. Here's the correct way:

# For a standalone factor column
numeric_column <- as.numeric(as.character(your_factor_column))

# In a tidyverse pipeline (if you're modifying a data frame)
your_data <- your_data %>%
  mutate(numeric_column = as.numeric(as.character(your_factor_column)))

Also, looking at your scraping code—you converted Firmname to a factor, but this isn't necessary for firm names (it can actually cause headaches later). You can safely remove the Firmname <- as.factor(Firmname) line.

3. Cleaning Up Your Web Scraping Code

Your scraping code works, but we can make it more concise and readable using tidyverse functions:

library(rvest)
library(tidyverse)

url <- "https://w2.brreg.no/kunngjoring/kombisok.jsp?datoFra=01.01.2019&datoTil=25.09.2019&id_region=100&id_fylke=-+-+-&id_niva1=2&id_bransje1=0"

nordland <- read_html(url) %>%
  # Extract and clean firm names and dates in one go
  tibble(
    `Firm name` = html_nodes(., "td td:nth-child(2) p") %>%
      html_text() %>%
      str_remove_all("\n| ") %>% # Remove newlines and spaces in one step
      .[-1], # Drop the first row (header)
    Date = html_nodes(., "td:nth-child(6) p") %>%
      html_text() %>%
      as.Date("%d.%m.%Y") %>%
      .[-1] # Drop the first row to match firm names
  ) %>%
  slice(1:1052) # Keep the first 1052 rows as you did originally

This avoids creating multiple intermediate data frames and uses str_remove_all() instead of repeated gsub() calls for cleaner code.

内容的提问来源于stack exchange,提问作者Kristian Bjerke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 07:44:15