You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何筛选满足全样本期特定条件的企业样本?

Solution to Filter Qualified Firms

Hey there! Let's work through this sample filtering problem together. Based on your dataset and the two criteria you've laid out, here's a clear, step-by-step solution using common data analysis tools:

Key Criteria Recap

First, let's restate the two requirements to make sure we're aligned:

  • Criterion 1: The firm's rating never changes across all observed years (e.g., Firm C is excluded because its rating shifts from 1 to 2)
  • Criterion 2: The firm has observations both before/including 2011 AND after 2011 (e.g., Firm B is excluded because it has no records post-2011)

Python (Pandas) Implementation

Step 1: Load and Prepare Test Data

First, let's recreate your sample dataset:

import pandas as pd

# Your raw data
data = {
    'firm': ['A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'],
    'year': [2010, 2011, 2012, 2010, 2011, 2010, 2011, 2012],
    'rating': [1, 1, 1, 1, 1, 1, 2, 2]
}
df = pd.DataFrame(data)

Step 2: Calculate Firm-Level Statistics

We'll group the data by firm to compute three key metrics:

  • Number of unique ratings (to check for rating changes)
  • Minimum observed year (to confirm pre-2011 records)
  • Maximum observed year (to confirm post-2011 records)
firm_summary = df.groupby('firm').agg(
    unique_ratings=('rating', 'nunique'),
    earliest_year=('year', 'min'),
    latest_year=('year', 'max')
).reset_index()

Step 3: Filter Qualified Firms

Now we apply our two criteria to select valid firms:

valid_firms = firm_summary[
    (firm_summary['unique_ratings'] == 1) &  # No rating changes
    (firm_summary['earliest_year'] <= 2011) &  # Has pre/2011 records
    (firm_summary['latest_year'] >= 2012)  # Has post-2011 records
]['firm']

Step 4: Get the Final Filtered Dataset

Use the valid firm list to subset your original data:

filtered_df = df[df['firm'].isin(valid_firms)]

The result will only include Firm A's records, which meet both criteria.


R (dplyr) Implementation

If you prefer R, here's an equivalent solution using the dplyr package:

Step 1: Load Data and Library

library(dplyr)

# Your raw data
data <- data.frame(
    firm = c('A', 'A', 'A', 'B', 'B', 'C', 'C', 'C'),
    year = c(2010, 2011, 2012, 2010, 2011, 2010, 2011, 2012),
    rating = c(1, 1, 1, 1, 1, 1, 2, 2)
)

Step 2: Summarize and Filter Firms

valid_firms <- data %>%
    group_by(firm) %>%
    summarize(
        unique_ratings = n_distinct(rating),
        earliest_year = min(year),
        latest_year = max(year)
    ) %>%
    filter(
        unique_ratings == 1,
        earliest_year <= 2011,
        latest_year >= 2012
    ) %>%
    pull(firm)

Step 3: Subset the Original Data

filtered_data <- data %>% filter(firm %in% valid_firms)

How This Works

  • By grouping at the firm level, we can easily check if a firm's rating stayed consistent (only 1 unique rating value)
  • Checking the earliest and latest years ensures the firm has presence both before/including 2011 and after 2011, which satisfies your second criterion

内容的提问来源于stack exchange,提问作者user8922408

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:20:48