基于Python/R的2018年1-5月限定条件海景房产均价预测技术问询
Got it, let's tackle this problem step by step. First, we need to clean up that unstructured "views" column to reliably flag all sea/beach/bay-facing properties. Then we'll filter down to the exact subset you need (200-300 sq.m., sea-view), and finally build a time-series forecast for Jan-May 2018.
Python 实现步骤
1. 数据清洗:识别海景房产
Those variant sea-view terms are a pain, but regex will save the day here. Just replace 'view_column' with your actual column name in the dataset:
import pandas as pd # 加载你的数据集(替换成你的文件路径) df = pd.read_csv('your_property_data.csv') # 定义匹配所有海景变体的正则表达式(忽略大小写) sea_view_pattern = r'(sea view|seaview|ocean view|water view|views to the sea|views of the sea|facing the sea|views to the beach|over the whole bay|over the bay)' # 添加标记列:1表示海景房产,0表示非海景 df['has_sea_view'] = df['view_column'].str.contains(sea_view_pattern, case=False, na=False).astype(int)
2. 筛选目标数据集
Now narrow down to properties between 200-300 sq.m. and with sea views, then aggregate historical data into monthly averages (since we're forecasting monthly prices):
# 替换成你的建筑面积列名、日期列名和价格列名 df_filtered = df[ (df['floor_area'] >= 200) & (df['floor_area'] <= 300) & (df['has_sea_view'] == 1) ].copy() # 按月份聚合均价,适配时间序列模型格式 df_monthly = df_filtered.groupby(pd.Grouper(key='date', freq='M'))['price'].mean().reset_index() df_monthly.columns = ['ds', 'y'] # Prophet要求的列名格式
3. 时间序列预测(用Prophet)
Prophet is perfect for monthly time-series forecasts—it handles seasonality automatically. We'll predict the average price for Jan-May 2018:
from prophet import Prophet # 初始化并拟合模型 model = Prophet(seasonality_mode='additive') model.fit(df_monthly) # 创建2018年1-5月的预测日期 future_dates = model.make_future_dataframe(periods=5, freq='M') future_dates = future_dates[future_dates['ds'].dt.year == 2018].head(5) # 生成预测结果 forecast = model.predict(future_dates) # 提取最终预测的均价 predicted_avg_prices = forecast[['ds', 'yhat']] print(predicted_avg_prices)
R 实现步骤
1. 数据清洗:识别海景房产
We'll use stringr for regex matching to catch all those sea-view variants. Replace view_column with your actual column name:
library(tidyverse) # 加载数据集 df <- read_csv("your_property_data.csv") # 定义海景正则模式(忽略大小写) sea_view_pattern <- regex( "sea view|seaview|ocean view|water view|views to the sea|views of the sea|facing the sea|views to the beach|over the whole bay|over the bay", ignore_case = TRUE ) # 添加标记列 df <- df %>% mutate(has_sea_view = as.integer(str_detect(view_column, sea_view_pattern)))
2. 筛选目标数据集
Filter for the 200-300 sq.m. range and sea views, then roll up historical data into monthly averages:
# 替换成你的建筑面积列名、日期列名和价格列名 df_filtered <- df %>% filter(floor_area >= 200, floor_area <= 300, has_sea_view == 1) # 按月份聚合均价,适配Prophet格式 df_monthly <- df_filtered %>% mutate(month = floor_date(date, "month")) %>% group_by(month) %>% summarize(avg_price = mean(price, na.rm = TRUE)) %>% rename(ds = month, y = avg_price)
3. 时间序列预测(用Prophet)
Use the R version of Prophet for consistent, easy-to-implement forecasts:
library(prophet) # 初始化并拟合模型 model <- prophet(df_monthly, seasonality.mode = "additive") # 创建2018年1-5月的预测日期 future_dates <- make_future_dataframe(model, periods = 5, freq = "month") future_dates <- future_dates %>% filter(year(ds) == 2018) %>% slice(1:5) # 生成预测结果 forecast <- predict(model, future_dates) # 提取最终预测的均价 predicted_avg_prices <- forecast %>% select(ds, yhat) print(predicted_avg_prices)
内容的提问来源于stack exchange,提问作者Bruno Benevides

