You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PHP的Boolean Retrieval(Binary Operation)实现及结果格式需求咨询

Boolean Retrieval Implementation in PHP for Your Query

Let's refine your existing code to handle the Boolean query "One AND Two AND NOT Three" and output results in the requested "Term Doc..." format. Here's a step-by-step solution tailored to your needs:

Step 1: Preprocess the Corpus

First, we'll clean up document processing to create a clear mapping of document IDs to their unique, lowercase terms — this makes retrieval faster and avoids case sensitivity issues.

Step 2: Parse & Map Query Terms

We’ll break down your query into terms and operators, then build a lookup table showing which documents contain each term.

Step 3: Execute Boolean Logic

We’ll compute the intersection of documents containing "One" and "Two", then exclude any documents that include "Three".

Full Working Code

<?php
$path = "Korpus";

// 1. Build a map of document IDs to their unique, lowercase terms
$docTerms = [];
foreach (glob("$path/Doc*.txt") as $file) {
    // Extract document ID from filename (e.g., Doc1.txt → 1)
    preg_match('/Doc(\d+)\.txt/', $file, $matches);
    $docId = $matches[1];
    
    // Process content: lowercase, split by spaces, remove duplicates/empty strings
    $content = strtolower(trim(file_get_contents($file)));
    $terms = array_unique(explode(" ", $content));
    $terms = array_filter($terms, fn($term) => !empty($term));
    
    $docTerms[$docId] = $terms;
}

// 2. Parse the Boolean query into terms and operators
$query = "One AND Two AND NOT Three";
$tokens = array_filter(explode(" ", strtolower($query)));

$terms = [];
foreach ($tokens as $token) {
    if (!in_array(strtoupper($token), ['AND', 'OR', 'NOT'])) {
        $terms[] = $token;
    }
}

// 3. Map each query term to the documents that contain it
$termToDocs = [];
foreach ($terms as $term) {
    $termToDocs[$term] = [];
    foreach ($docTerms as $docId => $docTermsList) {
        if (in_array($term, $docTermsList)) {
            $termToDocs[$term][] = $docId;
        }
    }
}

// 4. Execute the query logic: (One AND Two) AND NOT Three
$andResult = array_intersect($termToDocs['one'], $termToDocs['two']);
$finalResult = array_diff($andResult, $termToDocs['three']);

// 5. Output in the requested format
echo "Term\tDocuments\n";
echo "------------------------\n";
foreach ($termToDocs as $term => $docs) {
    echo ucfirst($term) . "\t" . (empty($docs) ? "None" : implode(", ", $docs)) . "\n";
}

echo "\nQuery: \"One AND Two AND NOT Three\"\n";
echo "Matching Documents: " . (empty($finalResult) ? "None" : implode(", ", $finalResult)) . "\n";
?>

Key Improvements & Notes

  • Case Insensitivity: All text is converted to lowercase so "One" and "one" are treated as the same term.
  • Cleaner Data: Duplicate terms and empty strings (from extra spaces) are filtered out to avoid retrieval errors.
  • Readable Output: First prints each term with its associated documents, then the final query result in a clear format.
  • Flexibility: If your corpus uses different filename patterns or needs punctuation removal, you can adjust the regex or add a step to strip special characters from content.

内容的提问来源于stack exchange,提问作者lowspec

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:37:50