PHP调用维基百科API提取infobox及短描述报错解决方法
报错原因
Array to string conversion提示:执行current($data['query']['pages'])后得到的是包含页面ID、标题、修订版本信息的多维数组,PHP不支持直接通过echo输出数组,因此触发提示,最终仅打印"Array"字符串。Undefined offset: 0提示:页面信息数组不存在索引为0的元素,接口返回的正文维基文本(wikitext)实际存储在$page['revisions'][0]['*']路径下,直接访问$data[0]属于读取不存在的数组键。
修正实现代码
<?php // 拉取页面0号章节原始内容 $url = "https://en.wikipedia.org/w/api.php?action=query&prop=revisions&rvprop=content&format=json&titles=twitter&rvsection=0"; $response = json_decode(file_get_contents($url), true); $page = current($response['query']['pages']); $wikitext = $page['revisions'][0]['*']; // 提取短描述:优先匹配官方短描述模板,无匹配时兜底取首句 $shortDescription = ''; if (preg_match('/\{\{Short description\|([^}]+)\}\}/i', $wikitext, $matches)) { $shortDescription = trim($matches[1]); } if (empty($shortDescription)) { $cleanText = preg_replace('/\{\{[^}]+\}\}/', '', $wikitext); $cleanText = strip_tags($cleanText); preg_match('/^[^.。!?]+[.。!?]/', trim($cleanText), $firstSentence); $shortDescription = trim($firstSentence[0] ?? ''); } // 提取Infobox完整维基文本,处理嵌套括号避免截断 $infoboxWikitext = ''; $infoboxStart = stripos($wikitext, '{{Infobox'); if ($infoboxStart !== false) { $braceCount = 0; $textLength = strlen($wikitext); for ($i = $infoboxStart; $i < $textLength; $i++) { if ($wikitext[$i] === '{' && isset($wikitext[$i+1]) && $wikitext[$i+1] === '{') { $braceCount++; $i++; } elseif ($wikitext[$i] === '}' && isset($wikitext[$i+1]) && $wikitext[$i+1] === '}') { $braceCount--; $i++; if ($braceCount === 0) { $infoboxWikitext = substr($wikitext, $infoboxStart, $i - $infoboxStart + 1); break; } } } } // 转换Infobox维基文本为HTML $data = ''; if (!empty($infoboxWikitext)) { $parseApi = "https://en.wikipedia.org/w/api.php?action=parse&text=" . urlencode($infoboxWikitext) . "&contentmodel=wikitext&format=json"; $parseResult = json_decode(file_get_contents($parseApi), true); $data = $parseResult['parse']['text']['*'] ?? ''; } // 验证结果 // var_dump($shortDescription); // 输出string(32) "American social networking service" // echo $data; // 输出完整Infobox对应的HTML代码 ?>
实现说明
- 短描述提取优先匹配维基官方的
{{Short description}}模板,准确率远高于直接截取文本首句,兜底逻辑适配未添加该模板的页面。 - Infobox提取采用括号计数逻辑,可正确处理Infobox内部嵌套的模板、标签,不会提前截断内容。
- 维基文本转HTML直接调用官方解析接口,无需自行实现复杂的wikitext解析规则,避免格式解析错误。
- 若需要优化请求效率,可调整接口参数一次性获取解析后内容,减少HTTP请求次数。
内容的提问来源于stack exchange,提问作者user19366702
相关产品推荐
相关产品推荐

