使用gsub删除议会协议PDF文本中的可变格式页脚
Got it, let's tackle this footer cleaning problem! The issue with your current gsub call is that it only targets the fixed starting text of the footer, but doesn't account for the variable page numbers (x of N) that come after. We need a regular expression that matches the entire footer block instead of just part of it.
Solution Code
Here's the adjusted code that will remove the full footer string:
string <- "this is the first page. 18th legislative period page 1 of 44 this is the second page. 18th legislative period page 2 of 44 and this is the third page" clean_string <- gsub("18th legislative period page \\d+ of \\d+", "", string) # Trim any extra whitespace left behind clean_string <- trimws(gsub("\\s+", " ", clean_string)) print(clean_string)
What This Does:
18th legislative period pagematches the fixed part of your footer.\\d+matches one or more digits (covers both the current page numberxand total pagesN).- The second
gsubandtrimwsclean up any extra spaces that might be left after removing the footer (so you don't get double spaces between sentences).
Output:
[1] "this is the first page. this is the second page. and this is the third page"
Bonus: Make It More Flexible
If you're dealing with multiple legislative periods (not just the 18th), you can adjust the regex to match any numbered legislative period:
clean_string <- gsub("\\d+th legislative period page \\d+ of \\d+", "", string)
This will work for 17th, 18th, 19th, etc.
内容的提问来源于stack exchange,提问作者Tobias Remschel
相关产品推荐
相关产品推荐

