使用PHP将JSON转CSV时如何提取HTML中的链接并去除其余标签内容
PHP JSON转CSV提取a标签href解决方案
你现有代码的问题在于strip_tags()函数会直接清除所有HTML标签及属性,无法保留<a>标签的href链接地址,只需替换递归处理JSON值的逻辑,用正则匹配提取所有a标签的href属性即可。
核心修改点
将原代码中递归处理内容的代码段:
array_walk_recursive($jsonDecoded, function (&$val) { $val = strip_tags($val); });
替换为以下逻辑:
array_walk_recursive($jsonDecoded, function (&$val) { // 匹配所有a标签的href属性 preg_match_all('/<a[^>]+href="([^"]+)"/i', $val, $matches); if (!empty($matches[1])) { // 多个链接用换行分隔,可按需改为逗号等其他分隔符 $val = implode(PHP_EOL, $matches[1]); } else { // 没有a标签的内容保留原纯文本,不需要可改为 $val = ''; 清空 $val = strip_tags($val); } });
修改后完整代码
<?php //products json $uri = 'http://my.site/RM_json'; $opts = array( 'http' => array( 'method'=>'GET', 'header'=>'Content-Type: application/octet-stream',) ); $context = stream_context_create($opts); $jsondata = file_get_contents($uri, false, $context); if($jsondata === false){ $error = error_get_last(); echo $error['message']; } $jsonDecoded = json_decode($jsondata, true); // 提取a标签href的处理逻辑 array_walk_recursive($jsonDecoded, function (&$val) { preg_match_all('/<a[^>]+href="([^"]+)"/i', $val, $matches); if (!empty($matches[1])) { $val = implode(PHP_EOL, $matches[1]); } else { $val = strip_tags($val); } }); $file = 'fileout.csv'; $fh = fopen($file, 'w'); fprintf($fh, chr(0xEF).chr(0xBB).chr(0xBF)); fputcsv($fh,["title", "field_risk_minimization_type","field_risk_1","field_hcpg","field_patient_card","field_healthcare_provider_checkl","field_dhpc"]); if (is_array($jsonDecoded)) { foreach ($jsonDecoded as $line) { if (is_array($line)) { fputcsv($fh,$line); } } } fclose($fh); header('Content-Description: File Transfer'); header('Content-Disposition: attachment; filename='.basename($file)); header('Expires: 0'); header('Cache-Control: must-revalidate'); header('Pragma: public'); header('Content-Length: ' . filesize($file)); header("Content-Type: application/vnd.ms-excel; charset=utf-8"); readfile($file); ?>
内容的提问来源于stack exchange,提问作者iappnet
相关产品推荐
相关产品推荐

