Active · OngoingHealth Technology Assessment · Medical Device Intelligence

Medical Device Intelligence Platform

A five-pipeline production platform that scales medical device horizon scanning, time-to-market intelligence, and FDA post-market AI surveillance from manual hundreds to evidence-grounded thousands — built for ECRI under an active FDA BAA contract.

ECRI · home.ecri.org

Scaling device intelligence beyond manual limits

ECRI's analysts were manually researching emerging medical devices — reading conference abstracts, regulatory filings, clinical-trial registries, and company websites — to answer two questions at scale: what new devices are coming, and how far are they from market at what regulatory risk class?

Doing this by hand capped coverage at a few hundred devices and made consistency impossible across different analysts and sessions. ECRI needed to cover thousands of devices with evidence-grounded answers, repeatably, and at a cost they could defend to stakeholders.

PredictIQ AI was brought in as an embedded AI engineering partner — owning the full lifecycle from R&D through production, working alongside ECRI's analysts week to week with multiple pipelines in production and new capabilities shipping on a roughly weekly cadence.

Client

ECRI

Global nonprofit health research organization and commercial intelligence arm

Users

Senior device analysts and clinical subject-matter experts producing horizon-scanning and time-to-market intelligence for health systems and payers

Engagement Type

Embedded AI engineering — full lifecycle ownership from R&D through production

Five Production Pipelines

A deliberate lab → production split: every pipeline is proven in the R&D lab before it ships. Production code mirrors the lab one-to-one; automated sync-checks police drift.

01

Time-to-Market

Estimates development phase and EU/IVDR risk class for curated device lists. Hybrid multi-model design: Gemini 2.5 Pro retrieves grounded evidence via Google Search (with caching), Claude Sonnet adjudicates and runs a market-status pre-flight before scoring.

Gemini 2.5 ProClaude SonnetGoogle GroundingEU/IVDR
02

Horizon Scanning

Discovers novel device candidates from seed lists plus automated monitoring of 72+ free sources — FDA clearances, RSS feeds, PubMed, EU MDR/IVDR registries, and HTA bodies. Produces a deduplicated dataset with a daily digest for analyst review.

72+ SourcesFDAPubMedEU MDRDaily Digests
03

Attribute Validation

Validates and corrects device attributes — name, company, class, markets, development stage — producing a confidence matrix that feeds back into the Time-to-Market pipeline. A single iteration lifted TTM coverage by double-digit percentage points.

Attribute ValidationConfidence MatrixFeedback Loop
04

EMBASE Novelty Mining

Multi-model LLM classification of conference abstracts against SME-authored protocols. Surfaces ranked novel-device candidates from published literature for analyst review.

LLM ClassificationConference AbstractsSME Protocols
05

FDA Post-Market AI Surveillance

Production monitoring for AI-enabled medical devices — ingesting hospital PSO data (RL6 XML), computing KPI baselines via Welford's online algorithm, and detecting performance degradation (sensitivity, specificity) with two-layer alerting: absolute thresholds plus statistical deviation from running baseline. Facility anonymization via opaque UUID tokens per PSWP legal requirements. Under active FDA BAA contract.

FDA BAAPSO DataStatistical MonitoringPSWP ComplianceBaseline Detection

Shared Device Catalog + FDA BAA Context

A canonical device registry with stable IDs and lifecycle status unifies identity across all five pipelines — eliminating duplicate records and enabling cross-pipeline tracking of the same device through its development arc.

Pipeline 05 (FDA Post-Market AI Surveillance) operates under an active FDA Business Associate Agreement — PSWP-protected hospital PSO data never leaves the secure processing perimeter; facility identity is anonymized via opaque UUID tokens before any downstream computation.

Embedded Engineering, Measured at Every Step

Human-in-the-Loop by Design

Analysts review daily digests and candidate lists. Their accept/reject decisions and SME overrides flow back into the pipelines. The AI is a force multiplier that experts steer — not a replacement for expert judgment.

Everything is Measured

Cost, token usage, coverage, and block-reason breakdowns are first-class metrics for every pipeline run. MLflow tracks experiment results; Metabase surfaces dashboards. Confidence thresholds only increase after SME accuracy validation.

Cost-Disciplined LLM Architecture

Aggressive MD5 caching keyed on model + prompt version means warm-cache reruns cost near zero. A full re-ground is a deliberate, budgeted decision — not a routine operation.

Honest About Limits

We diagnose coverage gaps as data problems or code problems — not model failures. The TTM coverage ceiling is a data limit (commercial-only evidence), so we steer toward new data sources rather than burning budget tuning a model that can't fix it.

Lab → Production Discipline

Every pipeline is built and iterated in a dedicated R&D lab repo before shipping. Production code mirrors the lab one-to-one; automated sync-checks police drift. The rule: prove it in the lab, then port it.

Measured Impact

10,422

Devices in the canonical catalog; 96.7% EMDN taxonomy code coverage across all five pipelines.

98%

Signal auto-rejection — 5,000+ weekly horizon-scanning signals reduced to ~100 for analyst review, with zero false negatives on known novel approvals.

~87%

Time-to-Market accuracy on SME-validated top-30 confidence sample (cardiac baseline, after 77 iterative improvement phases).

+12.1pp

TTM coverage lift from a single run of the attribute validation pipeline — from 46.4% to 58.5% across 3,189 devices.

90.1%

EMBASE novel device classification accuracy (Claude Sonnet, 817-record validation set); 100% novel recall on Carol Beckman's external review set.

Near-zero

Incremental cost for warm-cache reruns — MD5 caching keyed on model + prompt version; full re-ground is a deliberate, budgeted decision.

Tech Stack

Python 3.12Google Vertex AIGemini 2.5 Pro / FlashClaude Sonnet 4.6FastAPICelery + RedisPostgreSQLSQLAlchemy / AlembicMLflowMetabaseDockerGCP (Cloud Run / Cloud SQL)

Why this engagement defines our approach

The ECRI platform is a reference case for our core thesis: grounded, cost-instrumented, human-in-the-loop LLM systems that move cleanly from lab to production. It demonstrates the hybrid multi-model pattern, disciplined caching, and the research-mirrors-production architecture we bring to every engagement.

It also demonstrates what honest AI engineering looks like: when we hit a coverage ceiling, we diagnosed it as a data problem and steered the client toward new sources — not toward tuning a model that couldn't fix it.

Working on a similar problem?

If you're processing large volumes of domain-specific documents or building intelligence pipelines from unstructured sources, let's talk.