Megaparsec使用try时输入消耗异常,错误位置不符合预期
Hey, I’ve run into exactly this problem before when building parsers with Megaparsec! The root cause is almost certainly that your current setup fully parses the entire reserved word first before checking if it’s invalid—by the time you throw the error, the parser has already consumed all characters of the reserved word, so the error points to the end instead of the start.
Here’s how to fix it properly:
The key is to catch the reserved word as soon as you confirm it’s a full reserved word, without letting the parser consume extra characters beyond what’s needed to identify it. This way, the error position stays anchored to the first character of the reserved word.
Step 1: Define your reserved words as a set
First, make a set of reserved words for quick lookups (this is optional but makes checks cleaner):
import qualified Data.Set as Set reservedWords :: Set.Set String reservedWords = Set.fromList ["if", "else", "let", "in"] -- Add your reserved words here
Step 2: Rewrite the identifier parser with proper reserved word checking
Instead of parsing the full identifier first, we’ll use try to attempt matching a reserved word (and confirm it’s not part of a longer identifier), then fail immediately if we find one. If no reserved word is matched, we proceed to parse a normal identifier:
import Text.Megaparsec import Text.Megaparsec.Char import Data.Void (Void) type Parser = Parsec Void String identifier :: Parser String identifier = do -- Try to match a complete reserved word (ensure it's not part of a longer identifier) try $ do reserved <- choice (map string reservedWords) -- Make sure the reserved word isn't followed by identifier characters (like letters/underscores) notFollowedBy (alphaNumChar <|> char '_') fail $ "reserved word '" ++ reserved ++ "' cannot be used as an identifier" -- Parse a valid identifier if no reserved word was found (:) <$> letterChar <*> many (alphaNumChar <|> char '_')
Why this works:
- The
trycombinator lets us attempt matching a reserved word without permanently consuming characters if the match fails (e.g., if we seeifxinstead ofif, the parser will backtrack and parse it as a normal identifier). notFollowedByensures we don’t flag partial matches—soifis caught as a reserved word, butifxis treated as a valid identifier.- Since we throw the error immediately after confirming the reserved word, the parser’s position stays at the start of the reserved word, so your error message will show the correct line and column (e.g.,
1:1forifinstead of1:3).
Test it out:
- Input
foo→ parses successfully as"foo" - Input
bar123→ parses successfully as"bar123" - Input
if→ throws error at position1:1with messagereserved word 'if' cannot be used as an identifier - Input
else→ throws error at position1:1as expected
内容的提问来源于stack exchange,提问作者suhdonghwi

