You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PHP遍历多文件夹JSON/HTML文件时数据重复问题求助

问题分析与解决方案

你的核心问题是循环内数据未正确隔离,导致所有表格行复用了最后一次循环的数据,即便文件夹名称和链接正确,也只是因为这两个值在循环内被直接赋值覆盖,而其他数据因引用或变量作用域问题没有对应到当前文件夹。

常见原因及解决步骤

  1. 变量作用域/引用问题
    如果你在循环外部定义了存储数据的变量(比如字典、对象),每次循环仅修改该变量的属性而非创建新实例,最终所有表格行都会指向同一个变量的引用,显示最后一次循环的数据。

    • 解决:在循环内部初始化当前文件夹的专属数据容器(比如每次循环新建一个字典/对象),确保每个文件夹的数据独立存储。
  2. 路径验证(排除隐性错误)
    虽然你认为路径正确,但仍建议在循环内打印当前处理的文件路径,确认每个路径都对应到不同的文件夹,排除路径拼接时的隐性错误(比如未正确处理路径分隔符、遗漏文件夹名称)。

  3. 无状态的数据提取逻辑
    如果提取JSON/HTML数据的代码依赖全局变量或缓存,会导致每次提取的都是旧数据。确保每次循环都基于当前文件夹的路径重新读取、解析文件,不要复用之前的解析结果。

代码示例(Python版本)

import os
import json
from bs4 import BeautifulSoup

main_path = "./test-results"
# 过滤出主路径下的子文件夹(排除.和..)
folders = [f for f in os.listdir(main_path) if os.path.isdir(os.path.join(main_path, f))]
table_rows = []

for folder in folders:
    # 每次循环新建独立的数据容器
    current_row = {
        "folder_name": folder,
        "html_link": f"./{folder}/index.html",
        "stats": {},
        "page_title": ""
    }

    # 读取并解析statistics.json
    json_path = os.path.join(main_path, folder, "statistics.json")
    if os.path.exists(json_path):
        with open(json_path, "r", encoding="utf-8") as f:
            current_row["stats"] = json.load(f)

    # 提取index.html中的数据(以页面标题为例)
    html_path = os.path.join(main_path, folder, "index.html")
    if os.path.exists(html_path):
        with open(html_path, "r", encoding="utf-8") as f:
            soup = BeautifulSoup(f.read(), "html.parser")
            current_row["page_title"] = soup.title.string if soup.title else "无标题"

    # 将当前文件夹的数据加入行列表
    table_rows.append(current_row)

# 生成Markdown表格
print("| 文件夹名 | 测试报告链接 | 通过率 | 页面标题 |")
print("|----------|--------------|--------|----------|")
for row in table_rows:
    pass_rate = row["stats"].get("pass_rate", "N/A")
    print(f"| {row['folder_name']} | [查看报告]({row['html_link']}) | {pass_rate} | {row['page_title']} |")

代码示例(PHP版本)

$mainPath = "./test-results";
// 过滤有效文件夹
$folders = array_filter(scandir($mainPath), function($item) {
    return !in_array($item, ['.', '..']);
});

$tableRows = [];

foreach ($folders as $folder) {
    // 每次循环初始化当前文件夹的数据
    $currentData = [
        'folderName' => $folder,
        'htmlLink' => "./$folder/index.html",
        'stats' => [],
        'pageTitle' => ''
    ];

    // 读取JSON数据
    $jsonPath = rtrim($mainPath, '/') . "/$folder/statistics.json";
    if (file_exists($jsonPath)) {
        $currentData['stats'] = json_decode(file_get_contents($jsonPath), true) ?: [];
    }

    // 提取HTML标题
    $htmlPath = rtrim($mainPath, '/') . "/$folder/index.html";
    if (file_exists($htmlPath)) {
        $dom = new DOMDocument();
        @$dom->loadHTMLFile($htmlPath);
        $titleNode = $dom->getElementsByTagName('title')->item(0);
        $currentData['pageTitle'] = $titleNode ? $titleNode->nodeValue : '无标题';
    }

    $tableRows[] = $currentData;
}

// 渲染HTML表格
echo '<table border="1">';
echo '<tr><th>文件夹名</th><th>测试报告</th><th>通过率</th><th>页面标题</th></tr>';
foreach ($tableRows as $row) {
    $passRate = $row['stats']['pass_rate'] ?? 'N/A';
    echo "<tr>
            <td>{$row['folderName']}</td>
            <td><a href='{$row['htmlLink']}'>查看报告</a></td>
            <td>{$passRate}</td>
            <td>{$row['pageTitle']}</td>
          </tr>";
}
echo '</table>';

关键总结

  • 每次循环必须创建独立的数据容器,避免引用复用;
  • 确保文件路径在循环内动态拼接,打印路径验证正确性;
  • 数据提取逻辑要完全基于当前循环的文件夹路径,不依赖外部状态。

内容的提问来源于stack exchange,提问作者Amphata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:32:45