基于PHP的Boolean Retrieval(Binary Operation)实现及结果格式需求咨询
Boolean Retrieval Implementation in PHP for Your Query
Let's refine your existing code to handle the Boolean query "One AND Two AND NOT Three" and output results in the requested "Term Doc..." format. Here's a step-by-step solution tailored to your needs:
Step 1: Preprocess the Corpus
First, we'll clean up document processing to create a clear mapping of document IDs to their unique, lowercase terms — this makes retrieval faster and avoids case sensitivity issues.
Step 2: Parse & Map Query Terms
We’ll break down your query into terms and operators, then build a lookup table showing which documents contain each term.
Step 3: Execute Boolean Logic
We’ll compute the intersection of documents containing "One" and "Two", then exclude any documents that include "Three".
Full Working Code
<?php $path = "Korpus"; // 1. Build a map of document IDs to their unique, lowercase terms $docTerms = []; foreach (glob("$path/Doc*.txt") as $file) { // Extract document ID from filename (e.g., Doc1.txt → 1) preg_match('/Doc(\d+)\.txt/', $file, $matches); $docId = $matches[1]; // Process content: lowercase, split by spaces, remove duplicates/empty strings $content = strtolower(trim(file_get_contents($file))); $terms = array_unique(explode(" ", $content)); $terms = array_filter($terms, fn($term) => !empty($term)); $docTerms[$docId] = $terms; } // 2. Parse the Boolean query into terms and operators $query = "One AND Two AND NOT Three"; $tokens = array_filter(explode(" ", strtolower($query))); $terms = []; foreach ($tokens as $token) { if (!in_array(strtoupper($token), ['AND', 'OR', 'NOT'])) { $terms[] = $token; } } // 3. Map each query term to the documents that contain it $termToDocs = []; foreach ($terms as $term) { $termToDocs[$term] = []; foreach ($docTerms as $docId => $docTermsList) { if (in_array($term, $docTermsList)) { $termToDocs[$term][] = $docId; } } } // 4. Execute the query logic: (One AND Two) AND NOT Three $andResult = array_intersect($termToDocs['one'], $termToDocs['two']); $finalResult = array_diff($andResult, $termToDocs['three']); // 5. Output in the requested format echo "Term\tDocuments\n"; echo "------------------------\n"; foreach ($termToDocs as $term => $docs) { echo ucfirst($term) . "\t" . (empty($docs) ? "None" : implode(", ", $docs)) . "\n"; } echo "\nQuery: \"One AND Two AND NOT Three\"\n"; echo "Matching Documents: " . (empty($finalResult) ? "None" : implode(", ", $finalResult)) . "\n"; ?>
Key Improvements & Notes
- Case Insensitivity: All text is converted to lowercase so "One" and "one" are treated as the same term.
- Cleaner Data: Duplicate terms and empty strings (from extra spaces) are filtered out to avoid retrieval errors.
- Readable Output: First prints each term with its associated documents, then the final query result in a clear format.
- Flexibility: If your corpus uses different filename patterns or needs punctuation removal, you can adjust the regex or add a step to strip special characters from content.
内容的提问来源于stack exchange,提问作者lowspec
相关产品推荐
相关产品推荐

