如何在Jlex中通过正则表达式匹配独立的"String"单词?
Got it, let's sort out this regex problem so you only match standalone "String" words, not the ones embedded in longer identifiers like helloString or Stringhello.
The Problem with Your Current Regex
Your current (String) pattern matches any occurrence of the substring "String" anywhere in the text—even when it's part of a longer identifier. That's why it's picking up those unwanted matches.
The Solution
To target only standalone "String" words, we need to use word boundary anchors and adjust the rule order in your JLex code. Here's what to do:
- Use Word Boundaries: In JLex,
\<denotes the start of a word, and\>denotes the end of a word. This ensures "String" is surrounded by non-word characters (like spaces, line breaks, or the start/end of the input) instead of other identifier characters (letters, numbers, underscores). - Adjust Rule Order: Since "String" also fits the
IDpattern, we need to place the "String" rule before theIDrule. JLex uses longest-match first, but when matches are the same length, it picks the first rule that matches.
Modified JLex Code
Here's the updated code with the fixes highlighted:
import java.io.*; enum TokenType {Type_String,Identifier} class Token{ String text; TokenType type; Token(TokenType type,String text) { this.text=text; this.type=type; } public String toString() { return String.format("[%s,%s]",type,text); } } %% %class Lexer %public %function getNextToken %type Token %{ public static void main(String[] args) throws IOException { FileReader r = new FileReader("in.txt"); Lexer l = new Lexer(r); Token tok; while((tok=l.getNextToken())!=null){ System.out.println(tok); } r.close(); } %} %line %char SPACE=[\r\t\n\f\ ] ALPHA=[a-zA-Z] DIGIT=[0-9] ID=({ALPHA}|_)({ALPHA}|{DIGIT}|_)* %% // This rule now matches ONLY standalone "String" words <YYINITIAL> \<String\> {return new Token(TokenType.Type_String,yytext());} // ID rule comes second, so it won't override standalone "String" <YYINITIAL> {ID} {return new Token(TokenType.Identifier,yytext());} <YYINITIAL> {SPACE}* {} <YYINITIAL> . {System.out.println("error - "+yytext());}
How It Works
- The
\<String\>pattern ensures that "String" is a complete word:\<makes sure there's no identifier character (letter, number, underscore) immediately before "String"\>makes sure there's no identifier character immediately after "String"
- By placing the
\<String\>rule before theIDrule, we ensure that standalone "String" gets matched asType_Stringinstead ofIdentifier(since both would match the 6-character "String", and JLex picks the first matching rule for same-length matches).
Test Result
For your input String helloString String Stringhello helloStringhello, the output will be:
[Type_String,String] [Identifier,helloString] [Type_String,String] [Identifier,Stringhello] [Identifier,helloStringhello]
Which is exactly what you wanted—only the standalone "String" words are marked as Type_String.
内容的提问来源于stack exchange,提问作者Adham Mostafa

