基于R data.table按婚姻状态统计客户优惠券使用情况
Got it, let's build that summary table using data.table efficiently. Here's a step-by-step solution tailored to your needs:
First, we'll generate the sample data as you provided, then aggregate by marital status to get the metrics you want—total customers, coupon users, and the percentage of coupon users in each group. I'll also convert the 0/1 marital status values to readable labels for clarity.
Full Code
library(data.table) # Generate your sample dataset n <- 100000 DT <- data.table( customer_ID = 1:n, married = rbinom(n, 1, 0.4), coupon = rbinom(n, 1, 0.15) ) # Create the summary table summary_table <- DT[, .( Total_Customers = .N, # Count total customers per group Customers_using_Coupons = sum(coupon) # Sum coupon (0/1) to get count of users ), by = .(Marital_Status = ifelse(married == 1, "已婚", "未婚"))][, # Calculate percentage, rounded to 2 decimal places Percent_Using_Coupon := round(Customers_using_Coupons / Total_Customers * 100, 2) ] # Reorder columns to match your desired format setcolorder(summary_table, c("Marital_Status", "Customers_using_Coupons", "Total_Customers", "Percent_Using_Coupon")) # View the result summary_table
What This Does:
by = .(Marital_Status = ifelse(married == 1, "已婚", "未婚")): Groups the data by a readable marital status label instead of raw 0/1 values..N: data.table's shortcut for counting the number of rows (total customers) in each group.sum(coupon): Sincecouponis a binary 0/1 variable, summing it gives the total number of customers who used a coupon.- The
[, Percent_Using_Coupon := ...]line adds a new column with the percentage of coupon users, rounded for readability. setcolorderrearranges columns to match the order you mentioned.
Sample Output
When you run the code, you'll get a table that looks like this (values will vary slightly due to random data generation):
| Marital_Status | Customers_using_Coupons | Total_Customers | Percent_Using_Coupon |
|---|---|---|---|
| 未婚 | 9045 | 60089 | 15.05 |
| 已婚 | 5921 | 39911 | 14.83 |
That's it! This approach is efficient even for large datasets (like your 100k rows) since data.table operates in memory and uses optimized aggregation.
内容的提问来源于stack exchange,提问作者vijay29

