Python中如何在代码顶部指定数据集运行行数(含DataFrame与数组)
Great question! When working with large DataFrames or arrays, testing your logic on a smaller subset first saves tons of time—without touching your original dataset. Here are a few straightforward, clean approaches:
1. Control Loop Range with a Test Row Variable
Add a variable at the top of your script to define how many rows you want to test. Use min() to avoid errors if your dataset is smaller than your test size:
# Define test row count at the top—easily adjust here TEST_ROWS = 100 df_VisitorType_no = pd.DataFrame(columns=['VisitorType_no']) # Loop over min(test rows, total rows) to stay safe for i in range(min(TEST_ROWS, df.shape[0])): if df.loc[i,'VisitorType'] == 'Returning_Visitor': df_VisitorType_no.loc[i,'VisitorType_no'] = 1 elif df.loc[i,'VisitorType'] == 'New_Visitor': df_VisitorType_no.loc[i,'VisitorType_no'] = 2 else: df_VisitorType_no.loc[i,'VisitorType_no'] = 3
This way, you only need to tweak the TEST_ROWS value at the top—no changes to your core logic or original DataFrame required.
2. Create a Temporary Test Subset
For even clearer testing, make a separate temporary subset of your data (your original df stays untouched). This works for both DataFrames and NumPy arrays:
For Pandas DataFrames:
TEST_ROWS = 100 # Create a test subset—original df remains intact df_test = df.head(TEST_ROWS) df_VisitorType_no = pd.DataFrame(columns=['VisitorType_no']) for i in range(df_test.shape[0]): if df_test.loc[i,'VisitorType'] == 'Returning_Visitor': df_VisitorType_no.loc[i,'VisitorType_no'] = 1 elif df_test.loc[i,'VisitorType'] == 'New_Visitor': df_VisitorType_no.loc[i,'VisitorType_no'] = 2 else: df_VisitorType_no.loc[i,'VisitorType_no'] = 3
For NumPy Arrays:
If you're working with a NumPy array instead, the idea is identical:
import numpy as np TEST_ROWS = 100 # Create test subset of the array arr_test = arr[:TEST_ROWS] # Example processing logic for the test array result = np.zeros_like(arr_test) for i in range(len(arr_test)): if arr_test[i] == 'Returning_Visitor': result[i] = 1 # ... rest of your conditions
3. Bonus: Replace Loops with Pandas Vectorized Operations (More Efficient!)
Your original loop works, but Pandas is designed for vectorized operations which are way faster—especially for large datasets. You can rewrite your logic in one line, and still test it on a subset:
TEST_ROWS = 100 df_test = df.head(TEST_ROWS) # Map VisitorType to numeric values in one step df_test['VisitorType_no'] = df_test['VisitorType'].map({ 'Returning_Visitor': 1, 'New_Visitor': 2, 'Other': 3 # Assuming "else" refers to 'Other' }) # Verify the result first, then apply to full dataset if correct print(df_test[['VisitorType', 'VisitorType_no']].head()) # Once tested, apply to full DataFrame: # df['VisitorType_no'] = df['VisitorType'].map(...)
This approach avoids slow loops entirely, and testing on the subset is just as easy.
内容的提问来源于stack exchange,提问作者Leockl

