XProc中<p:filter>能否接收文档序列作为输入?
<p:filter> in XProc accept a sequence of documents as input? First, let's cut to the core of your question: <p:filter> is designed to operate on a single document at a time, and it does not natively support processing a sequence of documents directly—even if you set sequence="true" on its input port. Here's why your code is failing, and how to fix it:
Why your current code throws an error
The err:XD0006 error might feel misleading at first, but it points to a fundamental mismatch between what <p:filter> expects and what you're sending it.
Per the XProc 1.0 specification, <p:filter> behaves like an input port with a dynamic select expression—but this expression is evaluated against a single document node, not a sequence of documents. The default cardinality of <p:filter>'s source port is single document (sequence="false"), and even if you explicitly set sequence="true" in your step invocation, Calabash's implementation doesn't support applying the select expression across an entire document sequence.
Your test with <p:count> works because <p:count> is explicitly built to handle document sequences (it counts the number of documents in the input), which is a completely different use case than <p:filter>'s node-selection purpose.
The correct way to filter a document sequence
To filter each document in a sequence, you need to iterate over the sequence and apply <p:filter> to each document individually using <p:for-each>. This avoids the performance overhead of wrapping/unwrapping sequences and aligns with <p:filter>'s intended design.
Here's the revised code that will work as expected:
<?xml version="1.0" encoding="UTF-8"?> <p:declare-step xmlns:p="http://www.w3.org/ns/xproc" xmlns:c="http://www.w3.org/ns/xproc-step" version="1.0"> <p:input port="source" sequence="true"> <p:inline> <doc> <content>Hello world!</content> </doc> </p:inline> <p:inline> <doc> <content>Goodbye world!</content> </doc> </p:inline> </p:input> <p:output port="result" sequence="true"/> <!-- Iterate over each document in the input sequence --> <p:for-each> <p:input port="source" sequence="true"/> <!-- Apply filter to the individual document --> <p:filter select="//content"/> </p:for-each> </p:declare-step>
This code processes each input document separately, applies the //content selection to each, and outputs a sequence of the filtered <content> nodes (each as their own document node, matching your original output requirement).
Why the spec description might have confused you
The spec says <p:filter> behaves like an <p:input> with a select expression—and while <p:input> ports can accept sequences, the select expression on an input port is evaluated per document in the sequence. <p:filter> doesn't replicate this per-document evaluation automatically for sequences; you need to explicitly handle the iteration with <p:for-each>.
Alternative for cross-document filtering
If you need to filter the sequence based on conditions that span multiple documents (e.g., only keep documents that match a value present in another document), then wrapping the sequence into a single tree with <p:wrap-sequence> becomes necessary. But for simple per-document filtering, the <p:for-each> approach is far more efficient.
内容的提问来源于stack exchange,提问作者obdurodon

