Java基于正则表达式从文件行提取primaryKey和国家字段实现方案
Java文本字段提取实现
需求规则
- 从指定文本文件逐行读取内容,每行提取
primaryKey、3位大写国家代码country两个字段 - 单行样例文本:
bla,bla42bla()bla=bla+blablaprimaryKey="(ZAPDBHV7120D41A,USA,blablablablablabla - 匹配规则:
- 每行有且仅有1组目标字段,位置不固定,可出现在行首、行中、行尾
primaryKey无固定长度,仅由大写英文字母[A-Z]、数字[0-9]组成- 目标字段固定前缀为
primaryKey="(,前缀后拼接顺序为:primaryKey值 → 英文逗号 → 3位大写country值 → 英文逗号
输入输出样例
输入
bla,bla42bla()bla=bla+blablaprimaryKey="(ZAPDBHV7120D41A,USA,blablablablablabla bla++blabla()bla=bla+blablaprimaryKey="(AA45555DBMW711DD4100,ARG,bla [...]
期望输出
The primaryKey is ZAPDBHV7120D41A The country is USA The primaryKey is AA45555DBMW711DD4100 The country is ARG
完整实现代码
核心正则为primaryKey="\(([A-Z0-9]+),([A-Z]{3}),,通过两个捕获组分别提取primaryKey和country值。注意不要用Pattern.matches()(该方法要求整行完全匹配正则,无法匹配行内子串),改用Matcher.find()扫描行内符合规则的片段。
import java.io.BufferedReader; import java.io.FileReader; import java.util.regex.Matcher; import java.util.regex.Pattern; public class RegexExtract { // 预编译正则,避免循环内重复编译影响性能 private static final Pattern FIELD_PATTERN = Pattern.compile("primaryKey=\"\\(([A-Z0-9]+),([A-Z]{3}),"); public static void main(String[] args) throws Exception { final String targetFile = "my_file.txt"; try (BufferedReader br = new BufferedReader(new FileReader(targetFile))) { String line; while ((line = br.readLine()) != null) { Matcher matcher = FIELD_PATTERN.matcher(line); String primaryKey = ""; String country = ""; if (matcher.find()) { primaryKey = matcher.group(1); country = matcher.group(2); } System.out.printf("The primaryKey is %s%nThe country is %s%n%n", primaryKey, country); } } } }
代码说明:正则中
(属于特殊字符需要转义,Java字符串中转义字符需要用双反斜杠;将Pattern定义为静态常量预编译,是逐行大文件读取场景下的标准性能优化写法。
内容的提问来源于stack exchange,提问作者PatPanda
相关产品推荐
相关产品推荐

