You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中如何对DateType或Timestamp类型字段使用类似df = df.where("dateColumn <= 1950")的条件进行过滤?

Filtering Date/Timestamp Columns in PySpark: Why <= 1950 Won’t Work

Great question! Short answer: No, writing df.where("dateColumn <= 1950") won’t work as you expect for DateType or TimestampType columns in PySpark. Let me break down why, and show you the correct approaches.

The Problem with Using Integer 1950

When you pass 1950 as a literal in your filter condition, PySpark treats it as an integer—not a date. Since your dateColumn is a date/time type, comparing an integer to a date will either:

  • Throw a type mismatch error (if Spark’s type checker catches it), or
  • Silently return incorrect results (if it doesn’t).

The latter happens because Spark will interpret that integer as a count of days (for DateType) or milliseconds (for TimestampType) since the Unix epoch (1970-01-01). So 1950 would translate to around mid-1975—not 1950 at all!

Correct Ways to Filter Date/Timestamp Columns

Here are two reliable methods to filter your date column to values on or before 1950:

1. Using SQL String Syntax

Wrap your date in quotes (to treat it as a string) and use ISO date format, or explicitly convert the year to a date with to_date():

# Filter DateType to dates on or before 1950-12-31
df = df.where("dateColumn <= '1950-12-31'")

# If you only have the year, convert it to a full date using to_date()
df = df.where("dateColumn <= to_date('1950', 'yyyy')")

For TimestampType columns, use a full timestamp string:

df = df.where("timestampColumn <= '1950-12-31 23:59:59'")

2. Using PySpark Column API (Type-Safe Approach)

Use col() to reference your column, and either pass an ISO date string directly (PySpark auto-parses it) or use to_date() to convert the year:

from pyspark.sql.functions import col, to_date

# Filter DateType column
df = df.where(col("dateColumn") <= "1950-12-31")

# Convert year 1950 to a date and compare
df = df.where(col("dateColumn") <= to_date("1950", "yyyy"))

For TimestampType, use a timestamp string or to_timestamp():

from pyspark.sql.functions import to_timestamp

df = df.where(col("timestampColumn") <= to_timestamp("1950-12-31 23:59:59"))

内容的提问来源于stack exchange,提问作者JAdel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 08:59:08