为DCG添加解析约束:优化Graphic Token规则并保持双向性
Great question—handling exclusionary constraints in DCGs while keeping bidirectionality (so the grammar works both for parsing and generating tokens) can feel clunky if you're stuck with redundant checks, but there's a cleaner way to structure this.
First, let's recap the problem: we need to define graphic_token as one or more valid graphic characters (per ISO/IEC 13211-1:1995), but it must not start with the comment sequence /*. The original implementation splits the logic into single-character tokens and multi-character tokens (with a check to exclude /* as the first two characters), which works but repeats the graphic_token_char check unnecessarily.
Clean, Bidirectional Solution
We can refactor the DCG to handle the exclusion constraint upfront using a "negative lookahead" approach that's pure logic-friendly (so it stays bidirectional thanks to dif/2):
% Base definition for valid graphic token characters (matches original spec) graphic_token_char --> member("#$&*+-./:<=>?@^~\\ "). % Optimized graphic_token DCG graphic_token --> % Case 1: Single valid character (automatically avoids the /* prefix) graphic_token_char ; % Case 2: Multi-character token, explicitly avoiding /* as the start ( % Subcase 2a: First character is not '/' (so no risk of /*) [C], graphic_token_char, dif(C, '/') ; % Subcase 2b: First character is '/', but second is not '*' "/", [C], graphic_token_char, dif(C, '*') ), % Match any remaining valid characters kleene_star(graphic_token_char). % Helper DCGs (retained from original implementation) kleene_plus(NT) --> NT, kleene_star(NT). kleene_star(NT) --> "" | kleene_plus(NT). member(Xs) --> [X], { member(X, Xs) }.
Why This Works Better
- No redundant checks: Unlike the original code, we don't repeat the
graphic_token_charvalidation for the first two characters—we directly encode the valid starting patterns instead of excluding the invalid one after checking validity. - Bidirectional safety: Using
dif/2(instead of the one-way\=/2) ensures the grammar works both for parsing input strings into tokens and generating valid tokens from scratch.dif/2enforces logical inequality regardless of variable binding order, which is critical for bidirectional DCGs. - Readability: The logic is explicit: either we have a single character, or we start with a non-'/' character followed by more valid chars, or we start with '/' followed by a non-'*' character and more valid chars. There's no hidden redundancy to parse through.
Alternative: Post-Match Constraint (Less Efficient but Concise)
If you prefer a more declarative approach (at the cost of some potential backtracking), you can first match any valid sequence of graphic characters, then add a constraint to exclude those starting with /*:
graphic_token --> kleene_plus(graphic_token_char, Token), { % Either it's a single character, or the first two aren't /* (length(Token, 1) ; Token = [C1,C2|_], dif([C1,C2], "/*")) }. % Update helper DCGs to capture the full token kleene_plus(NT, [X|Xs]) --> NT([X]), kleene_star(NT, Xs). kleene_star(NT, []) --> "". kleene_star(NT, [X|Xs]) --> NT([X]), kleene_star(NT, Xs). graphic_token_char([X]) --> member("#$&*+-./:<=>?@^~\\ ", X). member(Xs, X) --> [X], { member(X, Xs) }.
This is more concise but can lead to unnecessary backtracking: it will first generate/parse a full token, then discard it if it starts with /*. For most use cases, the first (forward-checking) solution is better because it avoids wasted work.
内容的提问来源于stack exchange,提问作者repeat

