Java实现loadDocument方法提取link=后合法标识符的问题求助
Fixing the Link Identifier Extraction in Java
Looking at your code, there are a couple of key issues preventing it from correctly extracting and validating the identifiers after "link=":
- Incorrect line matching: Your current condition checks if the line matches
link="+name", which compares the line to a string combining "link=" with thenameparameter (not the identifier from the line itself). This is backwards—you need to extract the identifier from the line, not compare it to the method's input. - Broken validation regex: The
correctLinkmethod uses an invalid regex pattern that doesn't properly validate valid identifiers.
Here's the corrected code that addresses these issues and meets your requirements:
import java.util.Scanner; import java.util.regex.Matcher; import java.util.regex.Pattern; public class DocumentLoader { public static void loadDocument(String name, Scanner scan) { // Case-insensitive pattern to match lines starting with "link=" and capture the identifier Pattern linkPattern = Pattern.compile("^link=(.*)$", Pattern.CASE_INSENSITIVE); while (scan.hasNextLine()) { String line = scan.nextLine(); // Stop processing when we hit the "eod" line if (line.equals("eod")) { return; } // Check if the line matches the link pattern Matcher matcher = linkPattern.matcher(line); if (matcher.find()) { String identifier = matcher.group(1); // Validate the identifier and print if valid if (correctLink(identifier)) { System.out.println(identifier.toLowerCase()); } } } } // Validates if the input is a legal identifier: starts with letter, followed by letters/digits/underscores public static boolean correctLink(String link) { return link.matches("^[a-zA-Z][a-zA-Z0-9_]*$"); } // Example main method for testing public static void main(String[] args) { Scanner scanner = new Scanner(System.in); loadDocument("test", scanner); scanner.close(); } }
Key Changes Explained:
- Case-insensitive link matching: The
Pattern.CASE_INSENSITIVEflag ensures lines likeLINK=MyID,Link=abc123_, andlink=XYZare all detected correctly. - Identifier extraction: The regex
^link=(.*)$captures everything after "link=" into a group, which we retrieve withmatcher.group(1)to get the raw identifier. - Proper validation regex: The
correctLinkmethod now uses^[a-zA-Z][a-zA-Z0-9_]*$, which enforces:- The first character must be a letter (upper or lowercase)
- Subsequent characters can be letters, digits, or underscores (any number, including zero)
- Simplified flow: The loop processes each line, checks for the "eod" stop condition first, then handles link extraction and validation.
Example Test Cases:
- Input line:
link=abc→ Output:abc - Input line:
LINK=AbC_123→ Output:abc_123 - Input line:
link=123Invalid→ No output (starts with digit) - Input line:
link=_BadStart→ No output (starts with underscore) - Input line:
eod→ Stops processing
内容的提问来源于stack exchange,提问作者nme
相关产品推荐
相关产品推荐

