You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于字符串模式生成完整的分页URL

补全分页URL的R语言实现

现有如下R语言数据框,其中的URL分属两个不同主题(thee-t119532和thee-t143332),分页遵循s-0、s-10、s-20……的规律,但部分主题仅包含部分分页链接,需要补全所有缺失的分页,生成包含全部分页的完整数据框。

原始数据框

df <- data.frame(url = c("https://tapatalk.com/groups/thee-t119532-s0.html?sid=674002291c431ba23dd69c34e8a20217",
      "https://tapatalk.com/groups/thee-t119532-s50.html?sid=674002291c431ba23dd69c34e8a20217",
      "https://tapatalk.com/groups/thee-t119532-s60.html?sid=674002291c431ba23dd69c34e8a20217",
      "https://tapatalk.com/groups/thee-t119532-s70.html?sid=674002291c431ba23dd69c34e8a20217","https://tapatalk.com/groups/thee-t143332-s0.html?sid=674002291c431ba23dd69c34e8a20217",
      "https://tapatalk.com/groups/thee-t143332-s30.html?sid=674002291c431ba23dd69c34e8a20217"),
    page_id=c("0","50","60","70","0","30"))

目标完整数据框

df_full <- data.frame(url = c("https://tapatalk.com/groups/thee-t119532-s0.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s10.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s20.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s30.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s40.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s50.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s60.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t119532-s70.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t143332-s0.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t143332-s10.html?sid=674002291c431ba23dd69c34e8a20217",
"https://tapatalk.com/groups/thee-t143332-s20.html?sid=674002291c431ba23dd69c34e8a20217",
      "https://tapatalk.com/groups/thee-t143332-s30.html?sid=674002291c431ba23dd69c34e8a20217"),
page_id=c("0","10","20","30","40","50","60","70","0","10","20","30"))

解决方案

借助tidyverse工具包(dplyr处理数据,stringr处理字符串)完成,步骤如下:

  1. 加载依赖包
library(tidyverse)
  1. 拆分URL关键信息
    从原始URL中提取主题标识、URL固定前缀/后缀,同时转换分页ID为数值类型:
df_processed <- df %>%
  mutate(
    # 提取主题ID(如thee-t119532)
    topic_id = str_extract(url, "thee-t\\d+"),
    # 提取URL固定前缀(到主题ID后、分页前)和后缀(sid参数部分)
    url_prefix = str_replace(url, "(thee-t\\d+)-s\\d+\\.html(\\?sid=.*)", "\\1-s"),
    url_suffix = str_extract(url, "\\?sid=.*")
  ) %>%
  mutate(page_id = as.numeric(page_id))
  1. 生成全部分页并拼接URL
    按主题分组,生成从0到最大分页、步长为10的完整分页序列,再拼接成完整URL:
df_full <- df_processed %>%
  group_by(topic_id, url_prefix, url_suffix) %>%
  summarise(
    max_page = max(page_id),
    # 生成完整分页序列
    page_id = seq(from = 0, to = max_page, by = 10),
    .groups = "drop"
  ) %>%
  # 拼接完整URL
  mutate(
    url = str_c(url_prefix, page_id, ".html", url_suffix),
    page_id = as.character(page_id)
  ) %>%
  # 匹配目标数据框列顺序
  select(url, page_id)

执行后,df_full即为包含全部分页的完整数据框,与目标结果一致。

内容的提问来源于stack exchange,提问作者sd3184

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 07:45:49