You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中仅将美国选手数据行左移的实现方法

Fixing Right-Shifted Rows for US Athletes in Pandas DataFrame

Got it, let's break this down step by step. The core issue is that US athletes' birth city values (formatted as "City,State") got split during data import, shifting their entire row right by one column—so their birth country ended up in the rookie column, rookie in weight, and original weight in problem. We need to target only those US rows, shift them left, and clean up the columns to align with non-US athletes.

Step 1: Identify US Athlete Rows

First, we need a reliable way to flag which rows belong to US athletes. Two common approaches work here:

  • If only US rows have non-null values in the problem column (since non-US rows don't have the extra shifted data):
    import pandas as pd
    import numpy as np
    
    # Assume your DataFrame is named `df`
    us_mask = df['problem'].notna()
    
  • Or if the shifted rookie column for US rows contains "USA" (their original birth country):
    us_mask = df['rookie'] == 'USA'
    

Step 2: Shift US Rows Left by One Column

Now we'll target the shifted columns and move their values left to correct the alignment. The columns affected are birth city, birth country, rookie, weight, and problem:

# List of columns that need shifting
cols_to_shift = ['birth city', 'birth country', 'rookie', 'weight', 'problem']

# Shift values left: move each column's value to the previous column for US rows
df.loc[us_mask, cols_to_shift[:-1]] = df.loc[us_mask, cols_to_shift[1:]].values

# Clear the now-empty `problem` column for US rows
df.loc[us_mask, 'problem'] = np.nan

Step 3: Optional: Reconstruct the Full "City,State" Birth City

If you want to restore the original birth city format (combining the split city and state values), add this step:

# Merge the split city (now in `birth city`) and state (now in `birth country`)
df.loc[us_mask, 'birth city'] = df.loc[us_mask, 'birth city'] + ', ' + df.loc[us_mask, 'birth country']

# Set `birth country` back to "USA" for these rows
df.loc[us_mask, 'birth country'] = 'USA'

This approach is efficient even for thousands of rows because it uses vectorized operations (no slow loops) and only modifies the rows that need fixing.

内容的提问来源于stack exchange,提问作者JayBee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:58:39