如何使用Java基于预定义规则整理文本?附示例
下面针对两个常见的文本整理场景,给出具体的Java实现代码:
场景1:移除指定关键词之后的所有内容
需求:移除文本中"To help support the channel"及之后的所有内容(注:若只需移除关键词后的内容、保留关键词本身,可调整代码截取逻辑)
原始文本
About organic gardening, bees & beekeeping, blacksmithing, plant breeding, composting, natural farming, natural handicrafts, terra preta, biochar, compost tea, garden experiments, growing fungi & seed saving. To help support the channel: BTC Donation Address 14Tp6keFTi7YGXZdj1EfajwcZGTVQkkFfm
实现代码
public class TextCleaner { public static String cutAfterKeyword(String input, String keyword) { int keywordPos = input.indexOf(keyword); if (keywordPos != -1) { // 截取到关键词的起始位置,移除关键词及之后的内容 return input.substring(0, keywordPos).trim(); // 若需保留关键词,改为:return input.substring(0, keywordPos + keyword.length()).trim(); } // 未找到关键词时返回原文本 return input; } public static void main(String[] args) { String rawText = "About organic gardening, bees & beekeeping, blacksmithing, plant breeding, composting, natural farming, natural handicrafts, terra preta, biochar, compost tea, garden experiments, growing fungi & seed saving. To help support the channel: BTC Donation Address 14Tp6keFTi7YGXZdj1EfajwcZGTVQkkFfm"; String cleanedText = cutAfterKeyword(rawText, "To help support the channel"); System.out.println(cleanedText); } }
处理后效果
About organic gardening, bees & beekeeping, blacksmithing, plant breeding, composting, natural farming, natural handicrafts, terra preta, biochar, compost tea, garden experiments, growing fungi & seed saving.
场景2:保留About板块,移除Official Website及下方内容
需求:保留文本中About板块(包括后续的Links标题),移除"Official Website"及之后的所有内容
原始文本
About
Look no further if sports is your forte, head no further if you want to watch your favorite sport!
LinksOfficial Website
whatsapp.com/channel/0029Va6mIWNIyPtZGmLTup1wTwitter Link
twitter.com/StarSportsIndiaFacebook Link
facebook.com/starsportsindiaInstagram Link
instagram.com/starsportsindia
实现代码
用正则表达式匹配并截取目标内容,适合处理多行文本:
import java.util.regex.Matcher; import java.util.regex.Pattern; public class SectionFilter { public static String keepAboutAndLinks(String input) { // (?s) 开启单行模式,让.匹配换行符;(?=\\nOfficial Website) 正向预查,匹配到Official Website前的内容 Pattern pattern = Pattern.compile("(?s)^.*?(?=\\nOfficial Website)"); Matcher matcher = pattern.matcher(input); if (matcher.find()) { return matcher.group().trim(); } return input; } public static void main(String[] args) { String rawText = "About\n\nLook no further if sports is your forte, head no further if you want to watch your favorite sport!\nLinks\n\nOfficial Website\nwhatsapp.com/channel/0029Va6mIWNIyPtZGmLTup1w\n\nTwitter Link\ntwitter.com/StarSportsIndia\n\nFacebook Link\nfacebook.com/starsportsindia\n\nInstagram Link\ninstagram.com/starsportsindia"; String filteredText = keepAboutAndLinks(rawText); System.out.println(filteredText); } }
处理后效果
About
Look no further if sports is your forte, head no further if you want to watch your favorite sport!
Links
内容的提问来源于stack exchange,提问作者user352290

