列表值在for函数中的用法及SAS数据集变量YES响应占比计算报错求助
Hey there! Let's figure out why your code is spitting out 0 and get you the correct ratio you need.
What's Going Wrong With Your Original Code?
When you write for i in VS, you're actually looping through the column names of your DataFrame (like 'q1', 'q2', etc.), not the individual values in those columns. Since those column names are strings, they’ll never equal 1—so your yes count stays stuck at 0, hence the 0 result. Makes sense now, right?
Correct Ways to Calculate the YES (1) Response Ratio
Let's go through a few options, from the most efficient pandas-native approach to fixing your loop.
Option 1: Use Pandas Built-in Functions (Best Practice)
Pandas is made for this kind of calculation—no manual loops needed! It’s faster and cleaner:
import pandas as pd # Load your data (note the double backslash for Windows paths) lc = pd.read_sas("c:\\Downloads\\ps2_hedonic.sas7bdat", format='SAS7BDAT') Var = ['q'+str(z) for z in range(1,77)] VS = lc[Var] # Count total YES responses (1s) across all columns total_yes = (VS == 1).sum().sum() # Get total number of responses (all cells in the DataFrame) total_responses = VS.size # Calculate the ratio yes_ratio = total_yes / total_responses print(yes_ratio)
Here’s what’s happening:
(VS == 1)creates a boolean DataFrame whereTruemarks every 1 in your data..sum().sum()first sums up the 1s per column, then adds those column totals together for an overall YES count.VS.sizegives you the total number of values in the DataFrame (all responses combined).
Option 2: Fix Your For Loop
If you want to stick with a loop (maybe for learning purposes), you need to iterate through the actual values in the DataFrame, not just column names. Here’s how:
import pandas as pd lc = pd.read_sas("c:\\Downloads\\ps2_hedonic.sas7bdat", format='SAS7BDAT') Var = ['q'+str(z) for z in range(1,77)] VS = lc[Var] yes = 0 no = 0 # Loop through each row, then each value in the row for _, row in VS.iterrows(): for value in row: if value == 1: yes += 1 else: no += 1 yes_ratio = yes / (yes + no) print(yes_ratio)
Or, a more concise way to loop through all values directly:
# Alternative: Iterate through all flattened values for value in VS.stack(): yes += 1 if value == 1 else 0 no += 1 if value != 1 else 0
Using List Values in a For Loop
Your Var list is full of column names—here’s how to use that list to loop through each column’s values:
yes = 0 no = 0 # Loop through each column name in your Var list for col_name in Var: # Grab the column's data column = VS[col_name] # Loop through each value in the column for value in column: if value == 1: yes += 1 else: no += 1
This explicitly uses your Var list to target each column one by one.
Do You Need a While Loop?
Short answer: Nope. While loops are great for situations where you don’t know how many iterations you’ll need (like waiting for a user input or a condition to change). For iterating through structured data like a DataFrame, a for loop is way more readable and straightforward.
That said, if you’re curious how a while loop would work here (just for practice), here’s an example:
yes = 0 no = 0 current_row = 0 while current_row < len(VS): # Get the row at the current index row = VS.iloc[current_row] # Loop through values in the row for value in row: yes += 1 if value == 1 else 0 no += 1 if value != 1 else 0 # Move to the next row current_row += 1 yes_ratio = yes / (yes + no) print(yes_ratio)
But again, this is unnecessary when pandas has built-in tools for this!
内容的提问来源于stack exchange,提问作者Lyuhong Wang

