寻求个人程序开发建议:Python/Pandas CSV交互程序开发疑问
Hey Brandon, let's work through this together—you're already on a solid path using Pandas to handle your large CSV dataset, so let's fill in the missing pieces to get your program fully functional.
1. 确认方向:你的思路完全没问题
Using Pandas to load your master CSV and extract subsets for specific questions is absolutely the right call. Pandas makes slicing large datasets trivial, which would be way more cumbersome with just the basic csv module. Your current code to create dfq1 is a great start—we just need to build on it to complete the workflow.
2. 完善问题提示与带编号的列表展示
Right now, printing dfq1 directly will show the original index from your master CSV, not the user-friendly 1,2,3... numbering you want. Let's adjust that to display the list with sequential, intuitive numbers:
import pandas as pd # Load your master CSV (assuming column 0 is the original index, column 1 holds the string values) df = pd.read_csv("file.csv", encoding='unicode_escape', index_col=0) # Extract subset for Question 1 dfq1 = df.iloc[[0,18,25,33], 1] # Define your question prompt question = "Question 1: What type of issue are you reporting?" print(question) # Display options with user-friendly numbering (starts at 1) for idx, value in enumerate(dfq1, start=1): print(f"{idx} = {value}")
This loop uses enumerate() to add sequential numbers starting at 1, so users see exactly the format you described (e.g., 1 = Broken, 2 = Leaking).
3. 接收并验证用户输入
Next, we need to get the user's input and make sure it's a valid number that matches one of the options. We'll add basic error handling to avoid crashes from invalid inputs:
# Get user input with validation while True: try: user_input = int(input("Enter the number corresponding to your choice: ")) # Check if input falls within the valid range of options if 1 <= user_input <= len(dfq1): break else: print(f"Oops, please enter a number between 1 and {len(dfq1)}.") except ValueError: print("Invalid input—please enter a whole number.") # Map the user's 1-based input to the corresponding value and original index selected_value = dfq1.iloc[user_input - 1] original_master_index = dfq1.index[user_input - 1] # Optional, if you need to track the original CSV row
4. 写入结果到CSV
Finally, we'll write the question, selected number, and string to a results CSV. Using Pandas here keeps things clean, and we can append to the file so multiple responses are saved over time:
# Create a DataFrame for the result result_data = { "Question": [question], "Selected Number": [user_input], "Selected Value": [selected_value], "Original Master Index": [original_master_index] # Optional field } result_df = pd.DataFrame(result_data) # Append to results CSV (creates file if it doesn't exist, skips headers on subsequent writes) result_df.to_csv("responses.csv", mode='a', header=not pd.io.common.file_exists("responses.csv"), index=False) print("Response saved successfully!")
整合完整可复用代码
Putting it all together, here's a full script that's easy to extend for multiple questions:
import pandas as pd def process_question(question_text, df_subset): # Display question and formatted options print(question_text) for idx, value in enumerate(df_subset, start=1): print(f"{idx} = {value}") # Get and validate user input while True: try: user_input = int(input("Enter your choice number: ")) if 1 <= user_input <= len(df_subset): break print(f"Please enter a number between 1 and {len(df_subset)}.") except ValueError: print("Invalid input—only numbers are allowed.") # Retrieve selected data selected_value = df_subset.iloc[user_input - 1] original_index = df_subset.index[user_input - 1] # Save to results CSV result_data = { "Question": [question_text], "Selected Number": [user_input], "Selected Value": [selected_value], "Original Master Index": [original_index] } result_df = pd.DataFrame(result_data) result_df.to_csv("responses.csv", mode='a', header=not pd.io.common.file_exists("responses.csv"), index=False) print("Response saved!\n") # Load master dataset df = pd.read_csv("file.csv", encoding='unicode_escape', index_col=0) # Define all questions and their corresponding subsets question_map = { "Question 1: What type of issue are you reporting?": df.iloc[[0,18,25,33], 1], # Add more questions here as needed, e.g.: # "Question 2: How severe is the issue?": df.iloc[[5,12,20], 1] } # Process each question in sequence for q_text, q_subset in question_map.items(): process_question(q_text, q_subset)
额外小建议
- Scalability: Using a dictionary to map questions to subsets (like
question_mapabove) makes it super easy to add more questions without repeating code. - Error Handling: We added basic input validation, but you could expand it further (e.g., allowing users to exit the program mid-flow, handling empty subsets, etc.).
- CSV Management: The
header=not pd.io.common.file_exists(...)line ensures headers are only written once when the file is first created, so you don't get duplicate headers with multiple responses.
You were already on the right track—just needed to add the input handling, number formatting, and result writing pieces. Let me know if you need help adapting this to more questions or edge cases!
内容的提问来源于stack exchange,提问作者Brandon Harrelson

