如何筛选满足全样本期特定条件的企业样本?
Solution to Filter Qualified Firms
Hey there! Let's work through this sample filtering problem together. Based on your dataset and the two criteria you've laid out, here's a clear, step-by-step solution using common data analysis tools:
Key Criteria Recap
First, let's restate the two requirements to make sure we're aligned:
- Criterion 1: The firm's rating never changes across all observed years (e.g., Firm C is excluded because its rating shifts from 1 to 2)
- Criterion 2: The firm has observations both before/including 2011 AND after 2011 (e.g., Firm B is excluded because it has no records post-2011)
Python (Pandas) Implementation
Step 1: Load and Prepare Test Data
First, let's recreate your sample dataset:
import pandas as pd # Your raw data data = { 'firm': ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'], 'year': [2010, 2011, 2012, 2010, 2011, 2010, 2011, 2012], 'rating': [1, 1, 1, 1, 1, 1, 2, 2] } df = pd.DataFrame(data)
Step 2: Calculate Firm-Level Statistics
We'll group the data by firm to compute three key metrics:
- Number of unique ratings (to check for rating changes)
- Minimum observed year (to confirm pre-2011 records)
- Maximum observed year (to confirm post-2011 records)
firm_summary = df.groupby('firm').agg( unique_ratings=('rating', 'nunique'), earliest_year=('year', 'min'), latest_year=('year', 'max') ).reset_index()
Step 3: Filter Qualified Firms
Now we apply our two criteria to select valid firms:
valid_firms = firm_summary[ (firm_summary['unique_ratings'] == 1) & # No rating changes (firm_summary['earliest_year'] <= 2011) & # Has pre/2011 records (firm_summary['latest_year'] >= 2012) # Has post-2011 records ]['firm']
Step 4: Get the Final Filtered Dataset
Use the valid firm list to subset your original data:
filtered_df = df[df['firm'].isin(valid_firms)]
The result will only include Firm A's records, which meet both criteria.
R (dplyr) Implementation
If you prefer R, here's an equivalent solution using the dplyr package:
Step 1: Load Data and Library
library(dplyr) # Your raw data data <- data.frame( firm = c('A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'), year = c(2010, 2011, 2012, 2010, 2011, 2010, 2011, 2012), rating = c(1, 1, 1, 1, 1, 1, 2, 2) )
Step 2: Summarize and Filter Firms
valid_firms <- data %>% group_by(firm) %>% summarize( unique_ratings = n_distinct(rating), earliest_year = min(year), latest_year = max(year) ) %>% filter( unique_ratings == 1, earliest_year <= 2011, latest_year >= 2012 ) %>% pull(firm)
Step 3: Subset the Original Data
filtered_data <- data %>% filter(firm %in% valid_firms)
How This Works
- By grouping at the firm level, we can easily check if a firm's rating stayed consistent (only 1 unique rating value)
- Checking the earliest and latest years ensures the firm has presence both before/including 2011 and after 2011, which satisfies your second criterion
内容的提问来源于stack exchange,提问作者user8922408
相关产品推荐
相关产品推荐

