Botpress机器人对接Word文档文本读取API开发技术问询
Hey there! Let's work through your two Botpress + Word document integration questions to get everything running smoothly:
1. API Structure: Where to put your functions & how to connect them
You absolutely can keep everything in a single app.js file—this works great for simpler projects where you don't have tons of complex logic to manage. That said, here's how to structure it effectively:
- Single file approach: Wrap your textract document-reading logic into a reusable function directly in
app.js, then create an Express (or whatever framework you're using) API route that calls this function. Botpress can then hit this route via HTTP requests from its action files. - Scalable approach (for future growth): If you think you'll add more document types, authentication, or other features later, split your code into modules. For example:
- Create a
services/documentHandler.jsfile to hold all your textract logic (reading files, splitting paragraphs, etc.) - Import this module into your main
app.jsand hook it up to your API routes. This keeps your code organized and easier to maintain.
- Create a
2. Fetching & Returning Specific Word Document Paragraphs
The key here is to first convert the raw text from your Word doc into a structured array of paragraphs, then pull the specific one you need. Here's how to do it:
Step 1: Process the Word Document into Paragraphs
Use textract to pull the full text, then split it into paragraphs (adjust the splitting rule based on how your doc is formatted—empty lines are a common separator):
const getParagraphsFromWordDoc = (filePath) => { return new Promise((resolve, reject) => { textract.fromFileWithPath(filePath, (err, fullText) => { if (err) reject(err); // Split by empty lines to get individual paragraphs, filter out empty entries const paragraphs = fullText.split(/\n\s*\n/).filter(para => para.trim() !== ''); resolve(paragraphs); }); }); };
Step 2: Build an API Endpoint to Fetch Specific Paragraphs
Add this route to your app.js to let Botpress request a specific paragraph by index:
const express = require('express'); const textract = require('textract'); const app = express(); // ... include the getParagraphsFromWordDoc function here ... app.get('/api/get-paragraph', async (req, res) => { try { const { filePath, paraIndex } = req.query; if (!filePath || paraIndex === undefined) { return res.status(400).json({ error: 'Missing filePath or paraIndex parameter' }); } const paragraphs = await getParagraphsFromWordDoc(filePath); const index = parseInt(paraIndex); if (index < 0 || index >= paragraphs.length) { return res.status(404).json({ error: 'Paragraph index out of range' }); } res.json({ content: paragraphs[index] }); } catch (err) { res.status(500).json({ error: err.message }); } }); app.listen(3000, () => console.log('Document API running on port 3000'));
Step 3: Call the API from Botpress Action
Create a Botpress action file (e.g., fetchWordParagraph.js) to call your API and store the result in the session:
module.exports = async (bp, event, args, { session }) => { try { const apiResponse = await bp.http.get('http://localhost:3000/api/get-paragraph', { params: { filePath: './your-document-path.docx', paraIndex: '1' // Replace with the index of the paragraph you want } }); // Store the paragraph content in session to use in your bot's reply session.targetParagraph = apiResponse.data.content; } catch (err) { session.targetParagraph = 'Sorry, I couldn\'t retrieve that content right now.'; console.error('Error fetching paragraph:', err); } };
Step 4: Use the Content in Botpress Replies
In your bot's response template, just reference the stored value:
{{session.targetParagraph}}
Bonus: Skip the Separate API (Lightweight Option)
If you don't want to run a separate API, you can call textract directly inside a Botpress action—no extra app.js needed:
const textract = require('textract'); module.exports = (bp, event, args, { session }) => { textract.fromFileWithPath('./your-document-path.docx', (err, fullText) => { if (err) { session.response = 'Oops, I had trouble reading the document.'; return; } const paragraphs = fullText.split(/\n\s*\n/).filter(para => para.trim() !== ''); // Return the 3rd paragraph (index 2) session.response = paragraphs[2] || 'That paragraph doesn\'t exist.'; }); };
Then use {{session.response}} in your bot's reply template like you already do.
内容的提问来源于stack exchange,提问作者AVu

