如何通过cron获取谷歌文档原始内容而非HTML格式页面
问题原因
你当前使用的是谷歌文档的编辑页面URL,该地址返回的是完整的网页编辑器内容,自然包含HTML、CSS等前端代码,并非文档本身的原始内容。
解决方案
将URL中的edit?usp=sharing后缀替换为谷歌文档的导出接口后缀即可获取纯原始内容:
- 如需纯文本内容,替换为
export?format=txt - 如需docx格式文档,替换为
export?format=docx - 如需pdf格式文档,替换为
export?format=pdf
修改后的代码示例
基础版(file_get_contents实现)
<?php // 替换为纯文本导出地址 $exportUrl = 'https://docs.google.com/document/d/1IPNaCKwzdWRVOq30saI_pIBLcL62g4_2zMmjK54yD3E/export?format=txt'; $content = file_get_contents($exportUrl); // 移除谷歌导出文本自带的UTF-8 BOM头 $content = preg_replace('/^\x{EF}\x{BB}\x{BF}/u', '', $content); file_put_contents('import.txt', $content); ?>
兼容版(cURL实现,适配file_get_contents被禁用的环境)
<?php $exportUrl = 'https://docs.google.com/document/d/1IPNaCKwzdWRVOq30saI_pIBLcL62g4_2zMmjK54yD3E/export?format=txt'; $ch = curl_init($exportUrl); curl_setopt($ch, CURLOPT_RETURNTRANSFER, true); curl_setopt($ch, CURLOPT_FOLLOWLOCATION, true); // 跟随谷歌导出的302跳转 $content = curl_exec($ch); curl_close($ch); // 移除UTF-8 BOM头 $content = preg_replace('/^\x{EF}\x{BB}\x{BF}/u', '', $content); file_put_contents('import.txt', $content); ?>
注意事项
需要提前将目标谷歌文档的权限设置为「任何拥有链接的用户均可查看」,否则导出接口会返回403权限错误。
内容的提问来源于stack exchange,提问作者Dan
相关产品推荐
相关产品推荐

