如何使用Watson Knowledge Studio和Watson Discovery提取指定单位的值(基于指定报告的Matplot方法)
Alright, let’s walk through how to extract numerical values tied to specific units using Watson Knowledge Studio (WKS) and Watson Discovery, aligning with the Matplotlib-based approach from that empirical literature assessment report. Here’s a practical, step-by-step breakdown:
First, let’s map the report’s Matplotlib-focused workflow to Watson’s toolchain: the original study used pattern matching and validation to pull units like Cohen’s d, η², or p-values from literature. We’ll replicate that pattern recognition (plus add machine learning for edge cases) using WKS to define what counts as a "value + target unit" entity, then use Discovery to scale extraction across your document set.
Start here to teach the system to recognize your target numerical values and their associated units:
- Create a new WKS project tailored to academic literature (PDFs, text files from psychology/cognitive neuroscience papers work best).
- Define custom entity types (e.g.,
EffectSizeWithUnit,StatisticalValueWithUnit). For specificity, you can even use sub-types likeCohensDValueorPValue, but a flexible parent entity works too.- Manual annotation: Label 50-100 sample documents first—mark instances like
d = 0.85,η² = 0.12, orp < 0.05, linking the numerical value to its unit. This trains the ML model to pick up non-standard formatting (like.7instead of0.7orCohen’s dvsd). - Add rule-based extractors: Write regex patterns to cover consistent formats, just like the report’s pattern matching. Example patterns:
- For Cohen’s d:
(-?\d+\.?\d*) (Cohen's d|d) - For p-values:
p (?:=|<|>) (\d+\.?\d*|1e-\d+) - Wrap these in backticks in WKS’s rule editor to ensure proper parsing.
- For Cohen’s d:
- Manual annotation: Label 50-100 sample documents first—mark instances like
- Train and validate the combined ML+rule model: Test it against unannotated documents to adjust rules or add more labels if you miss edge cases.
Once your model is solid, move it to Discovery for bulk processing:
- Export the trained entity recognition model from WKS (look for the "Export model" option in the project settings).
- Import this model into your Watson Discovery instance, then enable it in your collection’s document processing pipeline. Make sure Discovery is set to parse PDFs/text correctly (enable OCR if you have scanned documents).
Now run the extraction to match the report’s bulk processing step:
- Upload your full set of literature documents to the Discovery collection.
- Use Discovery’s query interface (or API) to pull the entities you care about:
- Basic query to get all target entities:
enriched_text.entities.type:EffectSizeWithUnit - Filter by specific units:
enriched_text.entities.type:EffectSizeWithUnit AND enriched_text.entities.text:"Cohen's d"
- Basic query to get all target entities:
- Export the results as CSV or JSON—this gives you a clean list of values and their units, ready for analysis with Matplotlib (just like the original report did for statistical assessment and visualization).
- Cover edge cases in regex: Account for scientific notation (e.g.,
1.2e-5for p-values) and inconsistent spacing (e.g.,p<.001vsp < .001). - Add context filters: Use Discovery’s
passagesfeature to pull the paragraph around each entity—this lets you verify the value is actually an effect size/statistic, not a random number in the paper. - Iterate on annotations: If you notice the model missing values, go back to WKS and add those cases to your training set—this improves ML accuracy over time.
内容的提问来源于stack exchange,提问作者Blue_IoT_Mix

