Python新手求助:Pandas中DataFrame行与列切片及新表生成方法
Hey there! No worries at all—everyone starts somewhere with Python and Pandas. Let's break down what you need step by step, with clear examples you can follow.
First, let's assume you have your raw stock DataFrame ready (I'll include sample code to simulate it if you don't have real data yet). To create a new DataFrame with only the columns the user selects, you can directly index the original DataFrame with a list of column names.
Sample Raw DataFrame
import pandas as pd import numpy as np # Create simulated stock data with datetime timestamps dates = pd.date_range('2023-01-01', periods=10) stock_data = { 'timestamp': dates, 'open': np.random.uniform(100, 150, 10), 'close': np.random.uniform(100, 150, 10), 'high': np.random.uniform(100, 160, 10), 'low': np.random.uniform(90, 140, 10), 'volume': np.random.randint(100000, 500000, 10) } df = pd.DataFrame(stock_data)
Create New DataFrame with Selected Columns
Suppose the user wants timestamp, close, and volume—here's how to do it:
# Define the columns the user selects user_selected_cols = ['timestamp', 'close', 'volume'] # Option 1: Direct indexing (simple for column-only selection) new_df = df[user_selected_cols] # Option 2: Use .loc (more explicit, better for mixing row/column operations later) new_df = df.loc[:, user_selected_cols]
The : in .loc[:, user_selected_cols] means "select all rows"—we're just filtering the columns here.
Pandas has two main tools for slicing: .iloc (for position-based indexing, using integers) and .loc (for label-based indexing, using column/row names or conditions). Let's cover both.
Column Slicing
- Single column: You can use either
df['close'](works for all column names) ordf.close(only if the column name has no spaces/special characters). - Continuous columns: Use
.ilocwith integer ranges. For example, to select columns fromopentohigh(positions 1 to 3, 0-indexed):continuous_cols = df.iloc[:, 1:4] - Non-continuous columns: Use the same list-based approach as we did for the new DataFrame above.
Row Slicing
Position-based slicing with .iloc
Use integers to select rows by their position in the DataFrame:
# Select first 3 rows first_3_rows = df.iloc[:3] # Select rows 5 to 8 (0-indexed, so positions 4 to 7—left-inclusive, right-exclusive) rows_5_to_8 = df.iloc[4:8] # Select specific individual rows (e.g., 1st, 3rd, 5th rows) specific_rows = df.iloc[[0, 2, 4]]
Label/Condition-based slicing with .loc
This is more flexible for real-world use cases (like filtering by date or stock values):
# Filter rows where timestamp is on or after 2023-01-05 filtered_by_date = df.loc[df['timestamp'] >= '2023-01-05'] # Filter rows where close price > 120 AND volume > 300000 # Note: Use & for "and", | for "or", and wrap each condition in parentheses filtered_by_conditions = df.loc[(df['close'] > 120) & (df['volume'] > 300000)] # If you set timestamp as the DataFrame index (common for time-series data) df.set_index('timestamp', inplace=True) # Now you can slice directly by date range date_range_slice = df.loc['2023-01-02':'2023-01-06']
Mixed Row & Column Slicing
Combine row and column selection in one step:
# Select first 5 rows, only 'close' and 'volume' columns mixed_slice = df.iloc[:5, [2, 5]] # 2 = position of 'close', 5 = position of 'volume' # Or use .loc with conditions + column selection mixed_conditional = df.loc[df['high'] > 140, ['timestamp', 'high', 'low']]
内容的提问来源于stack exchange,提问作者Don Juan

