You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于列重组的卡方独立性检验实现需求

Solution: Restructure Data and Perform Chi-Square Test

Got it, let's walk through how to solve this problem step by step in R—super straightforward once we break it down:

1. Prep Your Data & Load Helpful Tools

First, we'll use the tidyr package to reshape our data easily (it's part of the tidyverse, so it's pretty standard for this kind of task). If you don't have it installed, grab it first:

install.packages("tidyr")
library(tidyr)

Next, let's turn your raw data into a proper R data frame:

# Create the original dataset from your input
raw_data <- data.frame(
  col1 = 1,
  col2 = 1,
  col3 = 3,
  col4 = 3,
  col5 = 5,
  Yes_Col_B = 7,
  No_Col_B = 9,
  Yes_Col_W = 3,
  No_Col_W = 2
)

# Pull out only the last four columns we care about
target_data <- raw_data[, c("Yes_Col_B", "No_Col_B", "Yes_Col_W", "No_Col_W")]

2. Reshape to Get "YesorNo" and "BorW" Columns

Right now our data is in wide format—we need to convert it to long format to split out our two variables. The pivot_longer function will do this perfectly by parsing the column names:

# Reshape the data and extract our two variables
reshaped_data <- target_data %>%
  pivot_longer(
    cols = everything(),
    names_to = c("YesorNo", "BorW"),
    names_pattern = "(Yes|No)_Col_(B|W)", # This regex splits the column names
    values_to = "Count"
  )

# Expand counts into individual observations (required for chi-square test)
# Since we have aggregated counts, we need to turn them into raw rows
final_data <- reshaped_data[rep(row.names(reshaped_data), reshaped_data$Count), ]
final_data$Count <- NULL # We don't need the count column anymore

3. Build Contingency Table & Run Chi-Square Test

Now we can create the contingency table and test if YesorNo depends on BorW:

# Create the observed contingency table
contingency_table <- table(final_data$YesorNo, final_data$BorW)
cat("Observed Contingency Table:\n")
print(contingency_table)

# Run the chi-square test for independence
test <- chisq.test(contingency_table)
cat("\nChi-Square Test Results:\n")
print(test)

# View the expected frequencies (the "chi-square table" you mentioned)
cat("\nExpected Frequencies Table:\n")
print(test$expected)

What This All Does:

  • The pivot_longer function splits our column names into two variables: YesorNo (grabs "Yes"/"No") and BorW (grabs "B"/"W"), while preserving the count values.
  • Expanding the counts into individual rows ensures the chi-square test works correctly (it expects raw observations, not aggregated counts).
  • The contingency table shows how many Yes/No responses we have for each B/W group.
  • The test output tells us if there's a significant relationship between the two variables. The expected frequencies table shows what counts we'd see if the variables were completely independent.

Example Output:

When you run the code, you'll get results like this:

Observed Contingency Table:
        BorW
YesorNo  B  W
    Yes  7  3
    No   9  2

Chi-Square Test Results:

	Pearson's Chi-squared test with Yates' continuity correction

data:  contingency_table
X-squared = 0.12698, df = 1, p-value = 0.7217

Expected Frequencies Table:
        BorW
YesorNo        B        W
    Yes  7.142857  2.857143
    No   8.857143  2.142857

In this case, the high p-value (0.7217) means we can't reject the null hypothesis—there's no significant evidence that YesorNo depends on BorW.

内容的提问来源于stack exchange,提问作者user35131

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 11:14:08