You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark中使用正则表达式移除URL列中的部分内容

PySpark处理Request URL列的实现方案

以下是针对Request URL列的几种常见处理场景的PySpark实现代码:

1. 提取URL路径部分

从包含协议、主机的URL中提取路径(如从https://example.com/api/v1/users?id=123提取/api/v1/users):

from pyspark.sql import functions as F

# 处理HTTP/HTTPS协议的URL
df = df.withColumn(
    "url_path",
    F.regexp_extract("Request URL", r"https?://[^/]+(/[^?]+)?", 1)
)

2. 提取指定查询参数值

提取URL中某个查询参数的内容(如提取id参数的值):

简单匹配(参数值为数字)

df = df.withColumn(
    "id_param",
    F.regexp_extract("Request URL", r"id=(\d+)", 1)
)

通用匹配(支持任意参数值,兼容参数在任意位置)

# 拆分查询参数列表,再过滤目标参数
df = df.withColumn(
    "query_params_list",
    F.split(F.regexp_extract("Request URL", r"\?(.*)", 1), "&")
).withColumn(
    "id_param_raw",
    F.expr("filter(query_params_list, x -> startsWith(x, 'id='))[0]")
).withColumn(
    "id_param",
    F.split("id_param_raw", "=")[1]
).drop("query_params_list", "id_param_raw")

3. 拆分URL各组成部分(Spark 3.0+)

Spark 3.0及以上版本提供了url_parse内置函数,可以直接解析URL的协议、主机、路径、查询等部分:

df = df.withColumn("parsed_url", F.url_parse("Request URL"))

# 提取各部分字段
df = df.withColumn("url_protocol", F.col("parsed_url.protocol"))
df = df.withColumn("url_host", F.col("parsed_url.host"))
df = df.withColumn("url_path", F.col("parsed_url.path"))
df = df.withColumn("url_query_string", F.col("parsed_url.query"))

4. 兼容异常URL的处理

针对格式不规范或为空的URL,添加空值判断避免报错:

df = df.withColumn(
    "url_path",
    F.when(
        F.col("Request URL").rlike(r"https?://"),
        F.regexp_extract("Request URL", r"https?://[^/]+(/[^?]+)?", 1)
    ).otherwise(F.lit(None))
)

内容的提问来源于stack exchange,提问作者Blue Clouds

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 14:47:13