SWI-Prolog中char_type/2返回大量字符的原因及解决方法
char_type/2 Results Hey there! Let's unpack what's going on with your SWI-Prolog queries and how to get the results you're expecting.
Why char_type(X, alnum) returns 816459 characters
SWI-Prolog's char_type/2 predicate is built on Unicode standards, not just ASCII. The alnum type matches every character in Unicode that's classified as a letter or number—this includes way more than just ASCII's A-Z, a-z, and 0-9. We're talking about every language's alphabet (Chinese characters, Japanese kana, Greek letters, etc.), plus all numeric symbols across Unicode. With Unicode containing hundreds of thousands of characters, that 816459 count is totally normal—it's just the total number of Unicode letters and numbers.
Why char_type(X, ascii) returns 2176 instead of 128 characters
Here's where things get a bit tricky with your Windows environment:
- SWI-Prolog's
asciitype behavior: While you'd expectchar_type(X, ascii)to only match the 127 standard ASCII characters (code points 0-127), on Windows with code page 852 (an Eastern European extended ASCII set) and your Prolog encoding set toascii, SWI-Prolog might be including extended ASCII characters (code points 128-255) in the results. - Duplicate results from
findall/3: The bigger issue is thatfindall/3collects every possible solution generated during backtracking. For some characters,char_type/2might generate multiple equivalent representations (like Unicode canonical equivalents or encoding-specific duplicates), leading to thousands of repeated entries in yourLasciilist.
When you run sort(Lascii, SortedLascii), the sort operation automatically removes duplicates and orders the characters—leaving you with the 128 standard ASCII characters you expected.
Fixes to get precise results
1. Get only ASCII alphanumeric characters
If you want just the ASCII letters and numbers (62 total), combine char_type/2 with code_type/2 to restrict to code points 0-127:
findall(X, (char_type(X, alnum), code_type(X, ascii)), AsciiAlnum), length(AsciiAlnum, N).
2. Get the full set of standard ASCII characters
To reliably get all 128 ASCII characters, use a query that directly targets code points 0-127:
findall(X, (between(0, 127, Code), char_code(X, Code)), Lascii), length(Lascii, N).
Alternatively, if you prefer using code_type/2, you can still use it with sort/2 to handle any duplicates:
findall(X, code_type(X, ascii), Lascii), sort(Lascii, SortedLascii), length(SortedLascii, N).
3. Verify character code points
If you want to confirm a character's ASCII code point, use char_code/2:
char_code('Z', Code). % Returns Code=90, which is the standard ASCII value
内容的提问来源于stack exchange,提问作者ktp

