You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jlex中通过正则表达式匹配独立的"String"单词?

Fixing the Independent "String" Matching in JLex

Got it, let's sort out this regex problem so you only match standalone "String" words, not the ones embedded in longer identifiers like helloString or Stringhello.

The Problem with Your Current Regex

Your current (String) pattern matches any occurrence of the substring "String" anywhere in the text—even when it's part of a longer identifier. That's why it's picking up those unwanted matches.

The Solution

To target only standalone "String" words, we need to use word boundary anchors and adjust the rule order in your JLex code. Here's what to do:

  1. Use Word Boundaries: In JLex, \< denotes the start of a word, and \> denotes the end of a word. This ensures "String" is surrounded by non-word characters (like spaces, line breaks, or the start/end of the input) instead of other identifier characters (letters, numbers, underscores).
  2. Adjust Rule Order: Since "String" also fits the ID pattern, we need to place the "String" rule before the ID rule. JLex uses longest-match first, but when matches are the same length, it picks the first rule that matches.

Modified JLex Code

Here's the updated code with the fixes highlighted:

import java.io.*;
enum TokenType {Type_String,Identifier}
class Token{
 String text;
 TokenType type;
 Token(TokenType type,String text) {
 this.text=text;
 this.type=type;
 }
 public String toString() {
 return String.format("[%s,%s]",type,text);
 }
}
%%
%class Lexer
%public
%function getNextToken
%type Token
%{
public static void main(String[] args) throws IOException {
 FileReader r = new FileReader("in.txt");
 Lexer l = new Lexer(r);
 Token tok;
 while((tok=l.getNextToken())!=null){
 System.out.println(tok);
 }
 r.close();
}
%}
%line
%char
SPACE=[\r\t\n\f\ ]
ALPHA=[a-zA-Z]
DIGIT=[0-9]
ID=({ALPHA}|_)({ALPHA}|{DIGIT}|_)*
%%
// This rule now matches ONLY standalone "String" words
<YYINITIAL> \<String\> {return new Token(TokenType.Type_String,yytext());}
// ID rule comes second, so it won't override standalone "String"
<YYINITIAL> {ID} {return new Token(TokenType.Identifier,yytext());}
<YYINITIAL> {SPACE}* {}
<YYINITIAL> . {System.out.println("error - "+yytext());}

How It Works

  • The \<String\> pattern ensures that "String" is a complete word:
    • \< makes sure there's no identifier character (letter, number, underscore) immediately before "String"
    • \> makes sure there's no identifier character immediately after "String"
  • By placing the \<String\> rule before the ID rule, we ensure that standalone "String" gets matched as Type_String instead of Identifier (since both would match the 6-character "String", and JLex picks the first matching rule for same-length matches).

Test Result

For your input String helloString String Stringhello helloStringhello, the output will be:

[Type_String,String]
[Identifier,helloString]
[Type_String,String]
[Identifier,Stringhello]
[Identifier,helloStringhello]

Which is exactly what you wanted—only the standalone "String" words are marked as Type_String.

内容的提问来源于stack exchange,提问作者Adham Mostafa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:26:35