在Pandas中依据strand列值修改start与end列值的技术咨询
Got it, let's solve this problem cleanly and efficiently. You want to tweak the start and end values only for rows where strand is '-', leaving '+' rows as-is. Here's how to do it right:
First, let's recreate your original DataFrame to work with:
import pandas as pd df = pd.DataFrame({ 'seqnames': ['A', 'B', 'C', 'D', 'E', 'F'], 'start': [50, 30, 10, 36, 40, 78], 'end': [453, 322, 432, 344, 321, 234], 'name': ['A']*6, 'number': [1,2,3,4,5,6], 'strand': ['-', '-', '-', '+', '+', '-'] })
Best approach: Use loc for efficient vectorized operations
This method is fast (even for large datasets) because it operates on entire columns at once instead of looping through rows:
# Update start column: add 1 where strand is '-' df.loc[df['strand'] == '-', 'start'] += 1 # Update end column: subtract 2 where strand is '-' df.loc[df['strand'] == '-', 'end'] -= 2
If you run print(df) now, you'll get exactly your expected output:
seqnames start end name number strand 0 A 51 451 A 1 - 1 B 31 320 A 2 - 2 C 11 430 A 3 - 3 D 36 344 A 4 + 4 E 40 321 A 5 + 5 F 79 232 A 6 -
Alternative: Using apply (for small datasets)
If you prefer a more explicit row-wise approach (good for small DataFrames, but slower on large ones), you can use apply:
df = df.assign( start=df.apply(lambda row: row['start'] + 1 if row['strand'] == '-' else row['start'], axis=1), end=df.apply(lambda row: row['end'] - 2 if row['strand'] == '-' else row['end'], axis=1) )
This will give you the same result, but keep in mind apply is less efficient than the vectorized loc method for big data.
内容的提问来源于stack exchange,提问作者chippycentra

