Medical Device Intelligence Platform
A five-pipeline production platform that scales medical device horizon scanning, time-to-market intelligence, and FDA post-market AI surveillance from manual hundreds to evidence-grounded thousands — built for ECRI under an active FDA BAA contract.
ECRI · home.ecri.org
Scaling device intelligence beyond manual limits
ECRI's analysts were manually researching emerging medical devices — reading conference abstracts, regulatory filings, clinical-trial registries, and company websites — to answer two questions at scale: what new devices are coming, and how far are they from market at what regulatory risk class?
Doing this by hand capped coverage at a few hundred devices and made consistency impossible across different analysts and sessions. ECRI needed to cover thousands of devices with evidence-grounded answers, repeatably, and at a cost they could defend to stakeholders.
PredictIQ AI was brought in as an embedded AI engineering partner — owning the full lifecycle from R&D through production, working alongside ECRI's analysts week to week with multiple pipelines in production and new capabilities shipping on a roughly weekly cadence.
Client
ECRI
Global nonprofit health research organization and commercial intelligence arm
Users
Senior device analysts and clinical subject-matter experts producing horizon-scanning and time-to-market intelligence for health systems and payers
Engagement Type
Embedded AI engineering — full lifecycle ownership from R&D through production
Five Production Pipelines
A deliberate lab → production split: every pipeline is proven in the R&D lab before it ships. Production code mirrors the lab one-to-one; automated sync-checks police drift.
Time-to-Market
Estimates development phase and EU/IVDR risk class for curated device lists. Hybrid multi-model design: Gemini 2.5 Pro retrieves grounded evidence via Google Search (with caching), Claude Sonnet adjudicates and runs a market-status pre-flight before scoring.
Horizon Scanning
Discovers novel device candidates from seed lists plus automated monitoring of 72+ free sources — FDA clearances, RSS feeds, PubMed, EU MDR/IVDR registries, and HTA bodies. Produces a deduplicated dataset with a daily digest for analyst review.
Attribute Validation
Validates and corrects device attributes — name, company, class, markets, development stage — producing a confidence matrix that feeds back into the Time-to-Market pipeline. A single iteration lifted TTM coverage by double-digit percentage points.
EMBASE Novelty Mining
Multi-model LLM classification of conference abstracts against SME-authored protocols. Surfaces ranked novel-device candidates from published literature for analyst review.
FDA Post-Market AI Surveillance
Production monitoring for AI-enabled medical devices — ingesting hospital PSO data (RL6 XML), computing KPI baselines via Welford's online algorithm, and detecting performance degradation (sensitivity, specificity) with two-layer alerting: absolute thresholds plus statistical deviation from running baseline. Facility anonymization via opaque UUID tokens per PSWP legal requirements. Under active FDA BAA contract.
Shared Device Catalog + FDA BAA Context
A canonical device registry with stable IDs and lifecycle status unifies identity across all five pipelines — eliminating duplicate records and enabling cross-pipeline tracking of the same device through its development arc.
Pipeline 05 (FDA Post-Market AI Surveillance) operates under an active FDA Business Associate Agreement — PSWP-protected hospital PSO data never leaves the secure processing perimeter; facility identity is anonymized via opaque UUID tokens before any downstream computation.
Embedded Engineering, Measured at Every Step
Human-in-the-Loop by Design
Analysts review daily digests and candidate lists. Their accept/reject decisions and SME overrides flow back into the pipelines. The AI is a force multiplier that experts steer — not a replacement for expert judgment.
Everything is Measured
Cost, token usage, coverage, and block-reason breakdowns are first-class metrics for every pipeline run. MLflow tracks experiment results; Metabase surfaces dashboards. Confidence thresholds only increase after SME accuracy validation.
Cost-Disciplined LLM Architecture
Aggressive MD5 caching keyed on model + prompt version means warm-cache reruns cost near zero. A full re-ground is a deliberate, budgeted decision — not a routine operation.
Honest About Limits
We diagnose coverage gaps as data problems or code problems — not model failures. The TTM coverage ceiling is a data limit (commercial-only evidence), so we steer toward new data sources rather than burning budget tuning a model that can't fix it.
Lab → Production Discipline
Every pipeline is built and iterated in a dedicated R&D lab repo before shipping. Production code mirrors the lab one-to-one; automated sync-checks police drift. The rule: prove it in the lab, then port it.
Measured Impact
Devices in the canonical catalog; 96.7% EMDN taxonomy code coverage across all five pipelines.
Signal auto-rejection — 5,000+ weekly horizon-scanning signals reduced to ~100 for analyst review, with zero false negatives on known novel approvals.
Time-to-Market accuracy on SME-validated top-30 confidence sample (cardiac baseline, after 77 iterative improvement phases).
TTM coverage lift from a single run of the attribute validation pipeline — from 46.4% to 58.5% across 3,189 devices.
EMBASE novel device classification accuracy (Claude Sonnet, 817-record validation set); 100% novel recall on Carol Beckman's external review set.
Incremental cost for warm-cache reruns — MD5 caching keyed on model + prompt version; full re-ground is a deliberate, budgeted decision.
Tech Stack
Why this engagement defines our approach
The ECRI platform is a reference case for our core thesis: grounded, cost-instrumented, human-in-the-loop LLM systems that move cleanly from lab to production. It demonstrates the hybrid multi-model pattern, disciplined caching, and the research-mirrors-production architecture we bring to every engagement.
It also demonstrates what honest AI engineering looks like: when we hit a coverage ceiling, we diagnosed it as a data problem and steered the client toward new sources — not toward tuning a model that couldn't fix it.
Working on a similar problem?
If you're processing large volumes of domain-specific documents or building intelligence pipelines from unstructured sources, let's talk.