PHP遍历多文件夹JSON/HTML文件时数据重复问题求助
问题分析与解决方案
你的核心问题是循环内数据未正确隔离,导致所有表格行复用了最后一次循环的数据,即便文件夹名称和链接正确,也只是因为这两个值在循环内被直接赋值覆盖,而其他数据因引用或变量作用域问题没有对应到当前文件夹。
常见原因及解决步骤
变量作用域/引用问题
如果你在循环外部定义了存储数据的变量(比如字典、对象),每次循环仅修改该变量的属性而非创建新实例,最终所有表格行都会指向同一个变量的引用,显示最后一次循环的数据。- 解决:在循环内部初始化当前文件夹的专属数据容器(比如每次循环新建一个字典/对象),确保每个文件夹的数据独立存储。
路径验证(排除隐性错误)
虽然你认为路径正确,但仍建议在循环内打印当前处理的文件路径,确认每个路径都对应到不同的文件夹,排除路径拼接时的隐性错误(比如未正确处理路径分隔符、遗漏文件夹名称)。无状态的数据提取逻辑
如果提取JSON/HTML数据的代码依赖全局变量或缓存,会导致每次提取的都是旧数据。确保每次循环都基于当前文件夹的路径重新读取、解析文件,不要复用之前的解析结果。
代码示例(Python版本)
import os import json from bs4 import BeautifulSoup main_path = "./test-results" # 过滤出主路径下的子文件夹(排除.和..) folders = [f for f in os.listdir(main_path) if os.path.isdir(os.path.join(main_path, f))] table_rows = [] for folder in folders: # 每次循环新建独立的数据容器 current_row = { "folder_name": folder, "html_link": f"./{folder}/index.html", "stats": {}, "page_title": "" } # 读取并解析statistics.json json_path = os.path.join(main_path, folder, "statistics.json") if os.path.exists(json_path): with open(json_path, "r", encoding="utf-8") as f: current_row["stats"] = json.load(f) # 提取index.html中的数据(以页面标题为例) html_path = os.path.join(main_path, folder, "index.html") if os.path.exists(html_path): with open(html_path, "r", encoding="utf-8") as f: soup = BeautifulSoup(f.read(), "html.parser") current_row["page_title"] = soup.title.string if soup.title else "无标题" # 将当前文件夹的数据加入行列表 table_rows.append(current_row) # 生成Markdown表格 print("| 文件夹名 | 测试报告链接 | 通过率 | 页面标题 |") print("|----------|--------------|--------|----------|") for row in table_rows: pass_rate = row["stats"].get("pass_rate", "N/A") print(f"| {row['folder_name']} | [查看报告]({row['html_link']}) | {pass_rate} | {row['page_title']} |")
代码示例(PHP版本)
$mainPath = "./test-results"; // 过滤有效文件夹 $folders = array_filter(scandir($mainPath), function($item) { return !in_array($item, ['.', '..']); }); $tableRows = []; foreach ($folders as $folder) { // 每次循环初始化当前文件夹的数据 $currentData = [ 'folderName' => $folder, 'htmlLink' => "./$folder/index.html", 'stats' => [], 'pageTitle' => '' ]; // 读取JSON数据 $jsonPath = rtrim($mainPath, '/') . "/$folder/statistics.json"; if (file_exists($jsonPath)) { $currentData['stats'] = json_decode(file_get_contents($jsonPath), true) ?: []; } // 提取HTML标题 $htmlPath = rtrim($mainPath, '/') . "/$folder/index.html"; if (file_exists($htmlPath)) { $dom = new DOMDocument(); @$dom->loadHTMLFile($htmlPath); $titleNode = $dom->getElementsByTagName('title')->item(0); $currentData['pageTitle'] = $titleNode ? $titleNode->nodeValue : '无标题'; } $tableRows[] = $currentData; } // 渲染HTML表格 echo '<table border="1">'; echo '<tr><th>文件夹名</th><th>测试报告</th><th>通过率</th><th>页面标题</th></tr>'; foreach ($tableRows as $row) { $passRate = $row['stats']['pass_rate'] ?? 'N/A'; echo "<tr> <td>{$row['folderName']}</td> <td><a href='{$row['htmlLink']}'>查看报告</a></td> <td>{$passRate}</td> <td>{$row['pageTitle']}</td> </tr>"; } echo '</table>';
关键总结
- 每次循环必须创建独立的数据容器,避免引用复用;
- 确保文件路径在循环内动态拼接,打印路径验证正确性;
- 数据提取逻辑要完全基于当前循环的文件夹路径,不依赖外部状态。
内容的提问来源于stack exchange,提问作者Amphata
相关产品推荐
相关产品推荐

