如何在新列提取每行的同比增长率并去除Bricks列中的数字?
Hey there! Let's tackle your two data manipulation questions one by one, using Python's pandas (the most common tool for this kind of task):
First, make sure your dataset has a year column and the metric column you want to calculate growth for (like sales, revenue, etc.). Here's how to compute YoY growth:
Basic YoY Calculation (Single Series)
If you're working with a single time series (no categories), sort your data by year first, then use pct_change() to get the growth rate (multiply by 100 to get a percentage):
import pandas as pd # Example dataframe df = pd.DataFrame({ "Year": [2020, 2021, 2022, 2023], "Sales": [2000, 2400, 3000, 3450] }) # Ensure data is ordered by year df = df.sort_values("Year") # Calculate YoY growth as a percentage df["YoY_Growth_Pct"] = df["Sales"].pct_change() * 100 # For non-percentage decimal format, remove the *100 # df["YoY_Growth"] = df["Sales"].pct_change() print(df)
Grouped YoY Calculation (Multiple Categories)
If you need to calculate YoY growth for different groups (e.g., different products, regions), use groupby() along with pct_change():
# Example grouped dataframe df = pd.DataFrame({ "Year": [2020, 2021, 2022, 2020, 2021, 2022], "Product": ["Brick", "Brick", "Brick", "Tile", "Tile", "Tile"], "Sales": [1500, 1800, 2250, 800, 920, 1050] }) # Sort by group and year first df = df.sort_values(["Product", "Year"]) # Calculate YoY growth per product df["YoY_Growth_Pct"] = df.groupby("Product")["Sales"].pct_change() * 100 print(df)
To strip all numeric characters from the Bricks column, use pandas' string method str.replace() with a regular expression that matches digits:
# Example dataframe with Bricks column df = pd.DataFrame({ "Bricks": ["ClayBrick_789", "RedBrick12", "450ConcreteBrick", "PlainBrick"] }) # Remove all digits (0-9) from the Bricks column df["Bricks_Cleaned"] = df["Bricks"].str.replace(r'\d+', '', regex=True) print(df)
The regex r'\d+' matches one or more consecutive digits, and replacing them with an empty string removes them entirely. If you need to handle edge cases (like digits separated by non-numeric characters), this regex still works since it targets all sequences of digits.
内容的提问来源于stack exchange,提问作者Sajjad Ali

