You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用gsub删除议会协议PDF文本中的可变格式页脚

Got it, let's tackle this footer cleaning problem! The issue with your current gsub call is that it only targets the fixed starting text of the footer, but doesn't account for the variable page numbers (x of N) that come after. We need a regular expression that matches the entire footer block instead of just part of it.

Solution Code

Here's the adjusted code that will remove the full footer string:

string <- "this is the first page. 18th legislative period page 1 of 44 this is the second page. 18th legislative period page 2 of 44 and this is the third page"
clean_string <- gsub("18th legislative period page \\d+ of \\d+", "", string)
# Trim any extra whitespace left behind
clean_string <- trimws(gsub("\\s+", " ", clean_string))
print(clean_string)

What This Does:

  • 18th legislative period page matches the fixed part of your footer.
  • \\d+ matches one or more digits (covers both the current page number x and total pages N).
  • The second gsub and trimws clean up any extra spaces that might be left after removing the footer (so you don't get double spaces between sentences).

Output:

[1] "this is the first page. this is the second page. and this is the third page"

Bonus: Make It More Flexible

If you're dealing with multiple legislative periods (not just the 18th), you can adjust the regex to match any numbered legislative period:

clean_string <- gsub("\\d+th legislative period page \\d+ of \\d+", "", string)

This will work for 17th, 18th, 19th, etc.

内容的提问来源于stack exchange,提问作者Tobias Remschel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:10:59