You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言按年份分组排序用户绘制分面直方图问题求助

Hey Gerry, let's work through this problem together! The error you're seeing happens because when you run df$user[order(df$year,df$spent)], you end up with a vector of 25 repeated usernames (since each user appears once per year), and factor levels must be unique. R throws an error because it hits the second occurrence of a username (like Laura at position 6) and can't use duplicate values for factor levels.

Here are two straightforward solutions to get your sorted, faceted plot working:


Solution 1: Use tidytext's reorder_within() (simplest for faceted sorting)

This function is built specifically for sorting categories within facets, so it handles the per-group ordering automatically.

First, install and load the required packages:

install.packages(c("tidyverse", "tidytext"))
library(tidyverse)
library(tidytext)

Then create your plot (note: we use geom_col() instead of geom_histogram() here because you're plotting individual user spending values, not distributions):

ggplot(df, aes(x = reorder_within(user, spent, year), y = spent)) +
  geom_col(fill = "#2c3e50") +
  facet_wrap(~year, scales = "free_x") + # Let each facet have its own x-axis
  scale_x_reordered() + # Clean up x-axis labels (removes auto-added year suffixes)
  labs(
    x = "User",
    y = "Amount Spent",
    title = "Annual Spending by User (Sorted by Amount)"
  ) +
  theme_minimal()

Solution 2: Manual grouping with dplyr and fct_reorder()

If you prefer not to use an extra package, you can group by year first to sort users within each year:

library(tidyverse)

# Create a sorted user factor within each year
df_sorted <- df %>%
  group_by(year) %>%
  mutate(user_sorted = fct_reorder(user, spent)) # Sort users by spent (ascending)
  # Use fct_reorder(user, -spent) if you want descending order

# Plot the sorted data
ggplot(df_sorted, aes(x = user_sorted, y = spent)) +
  geom_col(fill = "#2c3e50") +
  facet_wrap(~year, scales = "free_x") +
  labs(
    x = "User",
    y = "Amount Spent",
    title = "Annual Spending by User (Sorted by Amount)"
  ) +
  theme_minimal()

Why your original code failed:

When you ran df$user2=factor(df$user, levels = df$user[order(df$year,df$spent)]), the levels argument was a vector like c("Liz", "Laura", "Mark", "John", "Martha", "Mark", "Liz", ...) — full of duplicate usernames. Factor levels need to be unique, so R throws an error when it encounters the sixth element (the second "Mark") which is already in the levels list.

Either of the solutions above will fix this by handling the sorting per year group instead of trying to create a single global sorted factor.

内容的提问来源于stack exchange,提问作者Gerry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:09:43